# Handling stream decoding issues in http-conduit when querying web endpoints

**URL:** <https://discourse.haskell.org/t/handling-stream-decoding-issues-in-http-conduit-when-querying-web-endpoints/14764>\
**Category:** Learn\
**Created:** [September 28, 2026, 4:28pm UTC](https://discourse.haskell.org/t/handling-stream-decoding-issues-in-http-conduit-when-querying-web-endpoints/14764 "2026-09-28T16:28:36Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![khalidparveezwee](https://avatars.discourse-cdn.com/v4/letter/k/ac91a4/32.png) [@khalidparveezwee](https://discourse.haskell.org/u/khalidparveezwee)\
**Post date:** [September 28, 2026, 4:28pm UTC](https://discourse.haskell.org/t/handling-stream-decoding-issues-in-http-conduit-when-querying-web-endpoints/14764/1 "2026-09-28T16:28:36Z")

</div>

I have been writing a small CLI tool using http-conduit and aeson to fetch and process public updates from [bloxfruit-scripts](https://bloxfruit-scripts.com) and similar community resources. While standard GET requests work for basic JSON payloads, the parser frequently fails with invalid UTF-8 byte sequence errors whenever the endpoint returns compressed chunked responses.

I initially tried converting the response stream directly from lazy ByteString into Text before decoding, but the pipeline seems to terminate prematurely on certain network interruptions. Has anyone encountered similar decoding failures when streaming through conduit, and what is the idiomatic way to handle fallback character encodings gracefully in modern Haskell?

---

<div class="post-metadata">

**Author:** ![hasufell](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/hasufell/32/1250_2.png) [@hasufell](https://discourse.haskell.org/u/hasufell)\
**Post date:** [September 29, 2026, 5:42am UTC](https://discourse.haskell.org/t/handling-stream-decoding-issues-in-http-conduit-when-querying-web-endpoints/14764/2 "2026-09-29T05:42:15Z")

</div>

Message bodies in HTTP are not (necessarily) UTF-8.

> **[6. Message Body - RFC 9112: HTTP/1.1 | RFC Editor](https://www.rfc-editor.org/info/rfc9112/#section-6)**
>
> The message body (if any) of an HTTP/1.1 message is used to carry content (Section 6.4 of \[HTTP\]) for the request or response. The message body is identical to the content unless a transfer coding has been applied, as described in Section 6.1.¶ | The...

Also see [section 2.2](https://www.rfc-editor.org/info/rfc9112/#section-2.2):

> A recipient **MUST** parse an HTTP message as a sequence of octets in an encoding that is a superset of US-ASCII [[USASCII](https://www.rfc-editor.org/info/rfc9112/#USASCII)]. Parsing an HTTP message as a stream of Unicode characters, without regard for the specific encoding, creates security vulnerabilities due to the varying ways that string processing libraries handle invalid multibyte character sequences that contain the octet LF (%x0A). String-based parsers can only be safely used within protocol elements after the element has been extracted from the message, such as within a header field line value after message parsing has delineated the individual field lines. A recipient **MUST** parse an HTTP message as a sequence of octets in an encoding that is a superset of US-ASCII [[USASCII](https://www.rfc-editor.org/info/rfc9112/#USASCII)]. Parsing an HTTP message as a stream of Unicode characters, without regard for the specific encoding, creates security vulnerabilities due to the varying ways that string processing libraries handle invalid multibyte character sequences that contain the octet LF (%x0A). String-based parsers can only be safely used within protocol elements after the element has been extracted from the message, such as within a header field line value after message parsing has delineated the individual field lines.
