# DataFrame January 2026 Updates

**URL:** <https://discourse.haskell.org/t/dataframe-january-2026-updates/13512>\
**Category:** Announcements\
**Created:** [January 9, 2026, 6:55pm UTC](https://discourse.haskell.org/t/dataframe-january-2026-updates/13512 "2026-01-09T18:55:12Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![mchav](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/mchav/32/5143_2.png) [@mchav](https://discourse.haskell.org/u/mchav)\
**Post date:** [January 9, 2026, 6:55pm UTC](https://discourse.haskell.org/t/dataframe-january-2026-updates/13512/1 "2026-01-09T18:55:12Z")

</div>

### DataFrame update dump (0.3.1.1 → 0.4.0.5): faster, safer, and a lot more expressive

I’ve been heads-down shipping a _pile_ of improvements to **DataFrame** over the last few releases, and I wanted to share a “why you should care” summary (with some highlights + a couple examples).

#### Highlights

- **Ecosystem**

- **Performance wins across the board**

- **Decision trees**

- **Much nicer data cleaning + missing data ergonomics**

- **Expressions / schema evolution got smoother**

- **I/O + parsing hardening**

#### Example: expressive aggregation with conditionals (0.4.0.1)

```haskell
df
  |> D.groupBy [F.name ocean_proximity]
  |> D.aggregate
      [ "rand" .= F.sum (F.ifThenElse (ocean_proximity .== "ISLAND") 1 0)
      ]

```

#### Example: schema-friendly transformation pipeline (0.4.0.0+)

```haskell
print $ execFrameM df $ do
  is_expensive <- deriveM "is_expensive" (median_house_value .>= 500000)
  meanBedrooms <- inspectM (D.meanMaybe total_bedrooms)
  totalBedrooms <- imputeM total_bedrooms meanBedrooms
  filterWhereM (totalBedrooms .>= 200 .&& is_expensive)

```

If you’re doing ETL-y cleaning, feature engineering, quick stats, or want a Haskell-y dataframe that’s getting faster and more ergonomic every release: **this is a really good time to try the latest (0.4.0.4).**

Hoping to get a GSOC proposal for either Parquet writers or Arrow support so if you’d like to co-mentor please reach out.

---

<div class="post-metadata">

**Author:** ![mchav](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/mchav/32/5143_2.png) [@mchav](https://discourse.haskell.org/u/mchav)\
**Post date:** [January 9, 2026, 7:31pm UTC](https://discourse.haskell.org/t/dataframe-january-2026-updates/13512/2 "2026-01-09T19:31:57Z")

</div>

Btw also have a [PR](https://github.com/duckdblabs/db-benchmark/pull/144) to add dataframe to [Database-like ops benchmark](https://duckdblabs.github.io/db-benchmark/). After that’s in as a baseline (I think they said they’d release results after launching a new version of DuckDb) then I’ll probably need quite a bit of help with performance optimization.

---

<div class="post-metadata">

**Author:** ![romes](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/romes/32/2912_2.png) [@romes](https://discourse.haskell.org/u/romes)\
**Post date:** [January 9, 2026, 9:01pm UTC](https://discourse.haskell.org/t/dataframe-january-2026-updates/13512/3 "2026-01-09T21:01:55Z")

</div>

Fantastic work! Sounds like tons of excellent progress and ecosystem support is growing.

> probably need quite a bit of help with performance optimization.

Having a baseline which is easy to measure against will surely nerd snipe someone into trying to make it faster, myself included 😉
