# Summary
Haskell's `Fd`-level I/O API has these two key primitives:
* [`threa…dWaitRead :: Fd -> IO ()`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Control-Concurrent.html#v:threadWaitRead)
* [`threadWaitWrite :: Fd -> IO ()`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Control-Concurrent.html#v:threadWaitWrite)
It **is a program error** to have any outstanding I/O operations on an fd when that fd is closed. Equivalently, it is a program error to concurrently race closing an fd with these I/O operations on the fd. The fact that it is a program error is not adequately documented or emphasised.
The existing API has a kludge to partially cover up this program error:
* [`closeFdWith :: (Fd -> IO ()) -> Fd -> IO ()`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:closeFdWith)
This kludge is simulatenously [unreliable](#unreliable), [unnecessary](#unnecessary) and [unreasonably expensive](#unreasonably-expensive). Its existence [blocks the development](#blocking-progress) of new faster, more featureful and more scalable I/O managers in GHC.
The proposal ([detailed below](#proposal-and-migration)) is thus to deprecate and remove `closeFdWith`.
An [impact survey](#impact-survey) is included.
# Background to the problem
Module `Control.Concurrent` in base exports these two I/O concurrency "primitives":
* [`threadWaitRead :: Fd -> IO ()`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Control-Concurrent.html#v:threadWaitRead)
* [`threadWaitWrite :: Fd -> IO ()`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Control-Concurrent.html#v:threadWaitWrite)
Both operations document that:
> This will throw an [IOError](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Prelude.html#t:IOError) if the file descriptor was closed while this thread was blocked. To safely close a file descriptor that has been used with [threadWaitRead](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/Control-Concurrent.html#v:threadWaitRead), use [closeFdWith](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:closeFdWith).
Firstly, note that `closeFdWith` is _not_ exported from this module, or any primary API module in `base`. It is only exported from a `GHC.*` module. This is clearly unsatisfactory, it violates some API closure property.
The [docs for `closeFdWith`](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:closeFdWith) further clarify that the use of `closeFdWith` is mandatory
> Close a file descriptor in a concurrency-safe way as far as the runtime system is concerned (GHC only). If you are using [threadWaitRead](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:threadWaitRead) or [threadWaitWrite](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:threadWaitWrite) to perform blocking I/O, you _must_ use this function to close file descriptors, or blocked threads may not be woken.
(emphasis in the original)
Now, the `closeFdWith` docs also say
> Any threads that are blocked on the file descriptor via [threadWaitRead](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:threadWaitRead) or [threadWaitWrite](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:threadWaitWrite) will be unblocked by having IO exceptions thrown.
>
> Note that on systems that reuse file descriptors (such as Linux), using this function on a file descriptor while other threads can still potentially use it is always prone to race conditions without further synchronization.
>
> It is recommended to only call [closeFdWith](https://hackage-content.haskell.org/package/base-4.22.0.0/docs/GHC-Conc.html#v:closeFdWith) once no other threads can use the given file descriptor anymore.
It _does_ document that there's a race condition here, but does not emphasise it adequately. It is merely a recommendation to only close once the fd is no longer being used.
It is also worth noting that the documented behaviour of throwing exceptions to blocked threads is not universal: it is _only_ done by the threaded RTS I/O managers on Posix platforms. The fact that this behaviour is not universal shows that it is not something that programmers can rely on.
# History of `closeWithFd`
Kazu has done excellent [software archeologoy on this topic](https://gitlab.haskell.org/ghc/ghc/-/work_items/27801#note_699784). Rather than rephrase his work I will quote Kazu's research verbatim. Thus the "I" in this section is Kazu. Formatting and emphasis in the original.
---
I went through how `closeFdWith` came to exist. Three things seem worth putting on the record here.
**Its original justification was largely gone a day later.** `closeFd`, as it was then called, was added in [79cf73c5](https://gitlab.haskell.org/ghc/ghc/-/commit/79cf73c5089c6e59746e0e93d4d1c87d0d74cfa9) (2010-11-26) to fix [#4514](https://gitlab.haskell.org/ghc/ghc/-/work_items/4514). `threadWait` at that point looked like this:
```haskell
threadWait evt fd = do
m <- newEmptyMVar
Just mgr <- readIORef eventManager
_ <- registerFd mgr (\reg _ -> unregisterFd_ mgr reg >> putMVar m ()) fd evt
takeMVar m -- interrupted here, the registration stays behind
```
There is no cleanup on exception. The reproducer in [#4514](https://gitlab.haskell.org/ghc/ghc/-/work_items/4514) is a `timeout` interrupting `hGetLine`, so that is exactly what went wrong there, and the remedy chosen was to have the _closer_ sweep up rather than to have the interrupted thread clean up after itself.
The next day, [b9aeafa9](https://gitlab.haskell.org/ghc/ghc/-/commit/b9aeafa9c475b5d264af811eb8c7680765ec08ad) -- "Fix #4533 - unregister callbacks on exception, fixing a memory leak" -- added that cleanup, and [e30b2f4b](https://gitlab.haskell.org/ghc/ghc/-/commit/e30b2f4b675f8a39875afb9e702b6ccff905f879) gave it its present form:
```haskell
evt' <- takeMVar m `onException` unregisterFd_ mgr reg
```
The condition that made `closeFdWith` necessary was thus removed within a day of its introduction, for an unrelated reason, and whether it was still needed does not appear to have been revisited since.
**The `O(caps)` cost is not part of the design; it arrived two years later.** As introduced, it was `O(1)`:
```
closeFdWith close fd = do
Just mgr <- readIORef eventManager -- a single manager
M.closeFd mgr close fd
```
It became a sweep over every manager in [12f3fef5](https://gitlab.haskell.org/ghc/ghc/-/commit/12f3fef5ec52c1ec0958b674adcd981f48048428) (2012-12-22, "Parallel IO manager supports increasing and decreasing number of capabilities"), and [83ace021](https://gitlab.haskell.org/ghc/ghc/-/commit/83ace021a33fb6f4ce940f969467d6712cbaa858) (2021) then put that sweep under `uninterruptibleMask_`. The cost you are objecting to is the accumulation of those two later changes, not something the original design accepted.
A correction to my third point. I said that `closeFdWith` not being exported from Control.Concurrent was a gap. It is not an oversight: [dcc27a92](https://gitlab.haskell.org/ghc/ghc/-/commit/dcc27a92659328025a5bf1752ead31cb574e4160) (2010-12-06), "Drop closeFd from Control.Concurrent, rename to closeFdWith", removed it deliberately, and added the haddock sentence pointing the reader at `GHC.Conc.closeFdWith` in the same commit. It has been that way on purpose for sixteen years. If `closeFdWith` goes, the point is moot anyway.
# Background on Unix fds
Under Unix/Posix OSs, fds (file descriptors) are small integers used as pointers to kernel file objects. They are allocated to processes by the kernel. [POSIX.1-2017 (section 2.14)](https://pubs.opengroup.org/onlinepubs/9699919799/functions/V2_chap02.html) states:
> All functions that open one or more file descriptors shall, unless specified otherwise, atomically allocate the lowest numbered available (that is, not already open in the calling process) file descriptor at the time of each allocation.
This fd allocation policy makes it easy to use dense fd-indexed arrays, but it also means that fd numbers are:
1. reused;
2. reused concurrently (by other OS threads in the process); and
3. potentially reused very quickly.
Much like use-after-free bugs in memory unsafe languages, we must take this very seriously. Any (intended) use of a file (via an fd) after the fd pointing to the file has been closed may in fact be operating on an unrelated (recently opened) file.
# Unreliable
It is important to understand that a thread that is blocked in `threadWaitRead`/`threadWaitWrite` when another thread calls close on the same fd is actually in a _concurrent race_.
Thread A:
1. `threadWaitRead fd`
Thread B:
1. `close fd`
For it to be the case that thread A is blocked (on the same file) it means there is no "happens before" relation between step 1 of thread A and step 1 of thread B. A happens before relation here would mean that some synchronisation ensures the `close` only happens after the `threadWaitRead ` completes. Given the lack of such synchronisation we must say that these two steps in the two threads are executed _concurrently_. Indeed there is no synchronisation mechanism to guarantee that thread A is _actually_ blocked on the fd before the `close` happens. A perfectly possible schedule order could be that the close is executed first _before_ the `threadWaitRead`. In which case the wait is on an invalid or reused fd. This is a use-after-free bug.
In practice, timing can make it very likely that the wait happens before the close, but in princiuple this cannot be ensured. So in principle _every_ occurrence of a thread blocked in `threadWaitRead` when the fd is closed is a concurrent race, where one possible outcome of the race is to wait on an unrelated fd: a use-after-free bug.
Thus having `closeFdWith` automatically cancel any threads blocked on `threadWaitRead`/`threadWaitWrite` is _not a reliable solution_ to this race condition. Indeed it masks the bug by making it mostly benign in the cases that are easy to detect, but leaving the dangerous consequences in the hard to detect cases. So its behaviour is actively harmful.
# Unnecessary
A thread that calls `threadWaitRead`/`threadWaitWrite` on an fd will typically follow that up with a read or write on the same fd. Thus if the fd is closed while the thread is blocked in `threadWaitRead`/`threadWaitWrite` then if/when the thread wakes up the subsequent read/write will be either on an invalid fd or worse, a valid fd referring to an unrelated file.
One might reasonably guess that the original purpose of `closeFdWith` was to deal with this problem that it's unsafe to have threads blocked on `threadWaitRead` when the underlying fd is being closed. The history described above however shows that this is not the case. It was introduced to solve a proximate problem, which was solved shortly afterwards in a different way, but leaving `closeFdWith` in place.
The solution to the problem of racing `threadWaitRead`/`threadWaitWrite` against fd closing is _not to do that, at all, ever_. This must be part of the contract of `threadWaitRead`/`threadWaitWrite` (any any other existing or future APIs for I/O operations on fds).
It is one of the jobs of the higher level I/O layers (e.g. `Handle`, `Socket`) to deal with concurrency safely, and this includes ensuring races with close are handled safely.
One may also wonder if any of the platform APIs underlying `threadWaitRead`/`threadWaitWrite` require any post-deregristration steps that might necessitate a. The answer is no:
* Posix `poll()`/`select()`: these are inherently single-shot APIs
* Linux `epoll`: supports a single-shot event interest mode, and fds are removed automatically when the file is closed (after all fds referring to the file are closed)
* Linux `io_uring`: the fd polling support is single-shot by default
* FreeBSD/MacOS `kqueue`/`kevent`: supports a single-shot event interest mode.
* Win32: `IOCP`: requires `HANDLE` pre-association but no de-register step.
One might also reasonably wonder if the I/O manager itself needs a de-registration step to clean up its own data structures and prevent leaks. We note that it is _only_ the MIO I/O manager that does anything non-trivial in `closeFdWith`. Furthermore, we can note that the original MIO did not have `closeFdWith` at all, so it was not inherently needed for cleaning up resources (in the non-failing cases).
What would the behaviour be if `closeFdWith` is removed and threads _are_ still blocked on `threadWaitRead`/`threadWaitWrite` when the fd is closed?
* Posix `poll()`/`select()`: if the scheduler happens to poll for I/O readiness when the fd has been closed but not yet reused, then any thread blocked on `threadWaitRead`/`threadWaitWrite` will receive an exception (due to the use of a bad fd). If the fd is reused before the scheduler polls for I/O then the `threadWaitRead`/`threadWaitWrite` will wait on the same fd number that now referrs to a different file. This is obviously bad. It is also the _current behaviour_ in GHC in the non-threaded RTS. The I/O managers in the non-threaded RTS have never had the behaviour of throwing exceptions to threads when `closeFdWith` is called. For these I/O managers, `closeFdWith` just calls close.
* Linux `epoll`: epoll provides no notification when a fd is closed, so threads blocked on `threadWaitRead`/`threadWaitWrite` will block indefinitely. They _might_ get exceptions thrown later if the fd is reused since epoll does provide enough information to detect the reuse and thus the stale waiting threads.
* Linux `io_uring`: in this case a pending poll operation will keep the underlying file open, even when the fd is closed (the operation it self takes a reference to the underlying file). So not only would threads blocked on `threadWaitRead`/`threadWaitWrite` block indefinitely, but the file/socket would remain open. It is not clear that there's any efficient way to discover such stale waiters during the course of subsequent normal operations.
* FreeBSD/MacOS `kqueue`/`kevent`: the behaviour here is like epoll, from the man page: "Calling close(2) on a file descriptor will remove any kevents that reference the descriptor". There is no event delivered however so blocked threads will continue to block. I am not sufficiently familiar with this API to know if it is possible to discover stale waiters later if/when the fd is reused.
* Windows `IOCP`: all I/O operations are either completed or cancelled when a HANDLE is closed. This generates completion events on the IOCP, so it is possible to detect stale waiters and cancel them.
So the summary is that for all the stateful I/O APIs (everything but poll/select), we do not get the problem of `threadWaitRead`/`threadWaitWrite` completing successfully and leading to the use of a different fd by subsequent operations. On the other hand, it can potentially lead to threads hanging indefinitely, though there are some opportunities for detection. This is therfore somewhat easier to debug than trying to debug unexpected operations on a different file via a stale fd number.
# Unreasonably expensive
The performance problem with `closeFdWith` is that it does not and _can not_ scale with the number of capabilities/cores (i.e. +RTS -N).
The job of `closeFdWith` is to find out if there are _any_ outstanding Haskell threads on _any_ capability that are waiting on a particular fd. For a single capability this just requires a data structure to map fds to the threads waiting on them. It is not a local operation however: `closeFdWith` has to check this for _all capabilities_.
Consider these implementation strategies for I/O managers when using multiple capabilities/cores (i.e. +RTS -N):
1. A single I/O manager data structure shared across all capabilities. This would make finding all threads waiting on an fd cheap since it would be a single lookup in a single data structure. This obviously does not scale however and is not a reasonable design.
2. A sharded strategy with one I/O manager data structure per capability, but appropriate locking to allow cross-capability access when needed.
3. A share-nothing strategy with one I/O manager data structure per capability and no locking, so only the owning capability may access/modify the data structure. Any cross-capability actions must use message passing (or similar).
The current I/O manager used in GHC's threaded RTS follows the second strategy. The implementation of `closeFdWith` has to take a lock on the I/O tracking data structure for each of the capabilities. This is expensive: it is `O(caps)` and it reduces concurrency.
Future in-RTS I/O managers in GHC's threaded RTS will follow the third strategy. An implementation here would require passing a message to _every_ other capability so that they can search their I/O tracking data structures for threads waiting on a given fd and throw exceptions to them. And this must be synchronous to ensure that the exceptions are thrown before the file is closed, which would involve a second round of message passing and a new thread (TSO) blocking state for this purpose.
This is very bad for scalable network servers in Haskell. If we compare two hypothetical http servers, one that runs a separate process for each of `C` cores, and another that runs one process with `C` capabilities, and in each case we handle `N` connections per second per core. In the former design the aggregate cost of closing fds (sockets) is `O(C * N)` per second, since that's how many there are. While in the latter design the aggregate cost is `O(C^2 * N)`, because the cost of each close is `O(C)`.
And all this cost is for nothing -- in correctly written applications. In correctly written applications there are no outstanding operations on an fd when the fd is closed and thus all these searches will find nothing, but will still incurr the cost.
# Blocking progress
There are I/O managers in development now (based on platform APIs like `epoll`, `kqueue` and `io_uring`) that will follow a "shared nothing" design to maximise scalability on multiple cores and will only need to resort to cross-core message passing in rare cases. Keeping the cancellation behaviour would force us into "accidentally quadratic" `O(C^2 * N)` behaviour, reducing performance and limiting scalability.
# Proposal and migration
The important state to achieve is the state where `closeFdWith` does nothing more than simply call `closeFd` (i.e. actually close the fd).
The proposed steps are
1. properly document the contract of `threadWaitRead`/`threadWaitWrite`: that they **must not** be used concurrently with (or after) closing the fd;
2. remove `closeFdWith`'s behaviour of throwing exceptions to threads waiting on fds;
3. make `closeFdWith` unnecessary in any I/O manager implementation (e.g. for internal book keeping);
4. make `closeFdWith` do nothing more than `closeFd`;
5. deprecate `closeFdWith`;
6. eventually, remove `closeFdWith` entirely.
Optionally:
6. introduce some debugging function to help diagnose accidental races of `threadWaitRead`/`threadWaitWrite` with the closing of fds;
7. move `threadWaitRead`/`threadWaitWrite` to a more appropriate place in the libraries.
It may well be helpful to have an _expensive_ debugging feature to help find cases of accidental races: it would search all capabilities and inform the caller if any outstanding operations are still active (but not paper over the cracks by killing the offending threads).
It would be sensible to move all the `Fd` based operations from `Control.Concurrent` and put them somewhere else. These primitives are building blocks for I/O, networking and other system libraries, with non-trivial resource management constraints and should not be used directly in normal applications. Indeed, in 2010 [Simon Marlow opined](https://gitlab.haskell.org/ghc/ghc/-/work_items/4514#note_46664) that
> `threadWaitRead` and `threadWaitWrite` should not really be exported by `Control.Concurrent`, they're only there for historical reasons. They have a non-portable API (dependent on file descriptors). They also aren't very useful for most people, because they have to be used very carefully with non-blocking I/O operations.
The solution to all these fd lifetime and concurrency issues is higher level abstractions like `Handle` (or similar things at that level) that use locks and/or other representations than a raw fd that can ensure the safe use of fds. We should organise the I/O library functionality into layers to emphasise this point: that we have a low level (mostly portable) fd (or win32 HANDLE) layer with some pretty serious safety requirements, and then one or more higher level layers (like Handle and Socket) that provide a safer interface. We already sort-of have this, but the layering is not terribly clear.
# Impact survey
As one would expect, relatively few packages use `closeFdWith` directly.
The survey below is based on the Serokell hackage search:
https://hackage-search.serokell.io/?q=closeFdWith
I have categorised the packages based on their use.
### Full mutual exclusion
Several libraries follow a pattern of full mutual exclusion, using an `MVar` or equivalent to ensure I/O operations and close are not performed concurrently. The close operation also changes the state so that subsequent operations do not use the old fd. This fully ensures the proposed requirement that `threadWaitRead`/`threadWaitWrite` must not be used concurrently with (or after) closing the fd.
The following libraries use this pattern
* `ghc-internal-9.1401.0`: in the `Handle` implementation
* `linux-inotify-0.3.0.2`
* `posix-socket-0.3`
* `socket-0.8.3.0`
* `spacecookie-1.1.0.1`
### Atomic replacement with an invalid fd
The `network-3.2.9.0` package deserves special mention since it is so important.
It follows a pattern of atomic replacement with an invalid fd. This approach does not cancel or wait for ongoing operations, it just attempts to prevent new operations after a close. In this pattern the raw fd is kept in an `IORef` within an opaque type `Socket` and the API provide a `close` operation that atomically replaces the fd with -1. Only if the previous value was not -1 then it closes the fd. All other _subsequent_ operations will fail because they must retrieve the fd value, and if they use the -1 fd value then this will fail safely, since -1 is an alway-invalid fd number under unix.
This pattern is unfortunately flawed. It is subject to a race condition where an operation that is concurrent with close can still access a closed fd or a recycled fd. To pick a concrete example, consider racing thread A doing `sendBuf` with thread B doing socket `close`. Here is a possible interleaving:
1. Thread A: `fd <- readIORef ref`, getting a valid fd number (not -1)
2. Thread B: `atomicModifyIORef' ref` to atomically replace the fd with -1
3. Thread A: `c_send fd` 💥 use-after-free of the fd
The vulnerable window here is quite short. The most likely scenario where it would happen would be in heavily loaded multi-core network servers. The symptom would be random-looking corruption of data on freshly opened sockets, due to threads writing to unrelated recycled sockets. This is obviously a difficult scenario to debug.
The problem here is that there is no mutual exclusion. Grabbing a valid fd from the `IORef` does not ensure the fd remains valid for any period.
This problem already exists and is independent of the changes to `closeFdWith` proposed here. If this race is not fixed then the proposed changes to `closeFdWith` could add new bad behaviour: threads getting stuck in indefinitely blocking operations like network receive, whereas previously such threads would (often but not always) get exceptions thrown to them. This race could be fixed in such a way that the changes to `closeFdWith` are also not problematic.
### Re-export of or documentation that mentions `closeFdWith`
* `base-4.22.0.0`: re-exports `closeFdWith` and mentions in docs for `threadWaitRead`
* `base-io-access-0.4.0.0`: mentions `closeFdWith` in docs for `threadWaitRead'`
* `libssh2-0.2.0.9`: re-exports a wrapper around `threadWaitRead` and docs mention `closeFdWith`
* `planet-mitchell-0.1.0`: re-exports `closeFdWith`
### Relies on `closeFdWith`
* `hpqtypes-1.15.0.0`: `disconnect` uses `closeFdWith` and the comment indicates that it may rely on other threads being killed by this.
* `postgresql-libpq-0.11.0.0`: uses the same code pattern as `hpqtypes`
Further investigation may be needed here. Both libraries appear to give out their fds to the application to allow them to wait on the fd themselves. So perhaps in this sense they are not trying to abstract and ensure safe use of the fd. So this may just require clear documentation about the dangers of accidental use-after-free with raw fds.
### Already problematic (independent of `closeFdWith`)
* `cabal-install-bundle-1.18.0.2.1`: bundles a version of `Network.Socket` that does not provide any mutual exclusion of socket operations vs close, and is thus easily subject to use-after-close bugs.
### Other
* `haskell-names-0.9.9.1`: has the name, nothing else
*` ping-0.1.0.5`: explains in a comment why it does _not_ need to use `closeFdWith`