Buck2 build system for Cabal projects

This started as a little experiment and got a bit out of hand, but anyway: I’ve built an extension to Cabal to let you use Buck2 as the build system for a Cabal project. Why, might you ask? Well, it’s faster for one thing:

There are two parts to this:

The workflow is basically:

cabal buck2 --enable-tests
buck2 build //...  # builds everything
buck2 test //...  # runs all the tests

Instructions for how to use it and lots more details in the README for haskell-buck2.

I’d love to hear feedback and receive PRs. I would like to upstream this in Cabal, but the code needs cleaning up a fair bit first.

My not-so-secret evil plan is to build GHC this way, but we’ll see…

22 Likes

This sounds very exciting!!

1 Like

@simonmar I’m curious… after working on this, you must have noticed that doing any such things in cabal requires you to:

  1. identify all the obscure codepaths
  2. add new side effects to the existing blob of side effects (spitting out new files/artifacts/etc)

For a while I have been wondering if the entire cabal architecture could rather be done in a unix way:

  1. cabal would be split into many smaller binaries
  2. each of these would have a very distinct job (say… configuration vs dependency solving vs building vs installation, …)
  3. each of these binaries/phases would both consume well specified input and emit will specified output… effectively building a pipeline

In such an architecture the need for implementing things within cabal would be greatly reduced… you could simply replace the “build phase” with something entirely different as long as the output structure stays the same. The need for “hooks” would also be greatly reduced, as you can just add additional steps between two phases.

For the end user, all this could remain fairly invisible. The main cabal-install binary would essentially just be a plumbing wrapper.

I’m only superficially familiar with cabal on the code level, so my dreams could very well be shattered by reality. My guess is that it would essentially amount to a rewrite, so the RoI might be hard to argue for.

4 Likes

I wish Cabal/cabal-install were more modular, yes. In fact the first iteration of cabal buck2 was a separate tool that consumed the plan.json produced by cabal build all --only-dependencies. This was enough for a while, but eventually I needed to also configure all the local packages and process the results; building this into cabal-install was the only way. (I had Claude do most of the hard work so the fact that the internals are obscure just cost me tokens and I kept all my hair, at least what’s left of it). If there were a way to do the configure step and produce some well-defined machine-readable output that would be much better.

The problem that concerns me more is the extensibility (or lack of it) in a Cabal spec. The new hooks mechanism provides some extensibility, but… ugh. I don’t want to write my build rules using that API (or even in Haskell, if I’m honest). I’d much rather use a build system that’s designed for that purpose and has fast feedback and a good set of tooling and diagnostics.

My personal opinion is I don’t think Cabal should be a build system, and I’m dubious of adding more build system stuff to it. It was fine when the “build system” part of Cabal was basically just “call ghc --make”, but nowadays build systems need to do a LOT more, and it’s worth reconsidering whether building it into Cabal is the right way.

4 Likes

Pardon the naïveté here but since Cabal is meant to drive compilers plural, shouldn’t we on the contrary move the build system part into Cabal so that GHC/MHS/THC/MHC/Hazy do not have to reimplement this?

(Of course the modularity of Cabal is a very important topic, and I don’t think the project itself would be ready today for such a burden.)

PS: Did you by any chance come across this ticket already? Turn cabal-install into a build graph orchestrator · Issue #12121 · haskell/cabal · GitHub

1 Like

That’s disappointing to hear. My feeling is that cabal-install etc already has a lot of technical debt, and that deploying LLMs can only add to that. I understand that this project, right now, is mostly a prototype, but depending on LLMs like this makes me worried that it won’t scale beyond that.

2 Likes

That’s fair. To be honest I don’t think the code is terrible (but the comments often are!). I did push back a lot on the AI to avoid it doing tech debty things. But a better approach would be to refactor things to make the internals modular enough that adding something like cabal buck2 is just plugging together large pieces and some pretty-printing. That would be great, but it’s a much bigger job obviously.

2 Likes

I think that’s a reasonable question but I’m struggling to provide a reasonable answer. The design space is large, and not all designs result in duplication of effort and code, but some do. For example, if cabal buck2 can do all your building for you, then cabal build can just use cabal buck2 behind the scenes and you don’t need to implement build system stuff at all in Cabal. Obviously we’re not ready for that, but in that design you only have to support each compiler once. Another thing to bear in mind here is that implementing build support for a compiler might not be hard at all - MicroHs is pretty simple for example.

Thanks - I’m aware of the general move in this direction but I hadn’t seen that ticket specifically. Generally I think adding more build system stuff to Cabal doesn’t get us where we want to be long term. To support everything we eventually have to make Cabal into a fully-fledged general-purpose build system, and to do that is a lot of work.

3 Likes

I’m surprised by this. Most language ecosystems have a purpose-built build tool. Is that historical happenstance?
Edit : I need more coffee, I misread the quote.

i have a script that can help with that :grin: stfu.sh - Google-Search-AI-mode-coded script to remove all new comments from a git diff. Run it after your AI codes to shut it up. · GitHub

1 Like

My personal experience is probably unusual, but in both my personal projects and (two) places I’ve worked that use haskell, none of them have used cabal to build. The reasons are multiple languages, global cache, remote building, hermetic builds, “monadic” build graphs (eg generated source), custom scheduling, etc., all the usual build system bullet points. I already found ghc a bit closely tied to cabal, and I wouldn’t want it to get more tied. In one of those cases, I still want cabal as a version solver / dependency builder, and it can be done with some manual glueing via things like cabal build --only-deps and .ghc.environment… oh and I actually generate the cabal file from the shakefile, which is a bit circular. So, it’s not very clean. But cabal just can’t express many things I rely on. I also don’t (and wouldn’t want to!) declare .hs files manually in the .cabal file, and all its awkward copy pasting of deps lists in its non-extensible unique file format, I have more expressive ways.

So I’m with Simon on the idea of modularizing cabal and making the build backend swap-out-able, and nervous about ghc features that are sorted welded to cabal (backpack is the example I can think of, though it didn’t seem to catch on). ghc --make is actually pretty great for pure haskell since it does fancy recompile avoidance, but it’s a big not-pure-haskell world out there.

I don’t know enough about multiple haskell compilers and their differing needs to have an informed idea. But surely a real extensible build system could handle those things?

But, the world of “real extensible build systems” is also not easy, so I don’t have a recommendation. bazel is crazy complicated (I hear), I don’t know about buck2, shake is more of a construction kit, nix is an all-consuming lifestyle choice / religion. Maybe buck2 is The One?

3 Likes

I am very happy to see this @simonmar. I did some work on buck2/cabal interoperation, there is a lot of potential in this space (although sometimes hidden in the status-quo).

1 Like

i gotta be honest - it is The One. Between this work and Mercury’s, the community should lean into buck2 imo.

1 Like

you remember that pusher article about Haskell’s GC latency scaling with working set? and how they migrated to Go? and then compact regions and the non-moving GC eventually followed?

i think buck2 could be the equivalent vis a vis the Scarf migration away to Python for agentic reasons.

EDIT: I think a big thing here could be a pure Haskell implementation of buck2. à la hnix. That would really ease things.

Choosing your own build system for a standalone project is not a problem, and there are plenty to choose from. Buck2 is what I would choose today. But the real design challenge is how to do that and have a Cabal package that you can upload to Hackage and cabal install, without duplication or horrible hacks.

Design 1: Add support for build-type: Buck2 to Cabal. The experience could be made smooth: cabal build automatically invokes Buck2; you can customise your build as much as you like; Cabal would bundle haskell-buck2 somewhere under ~/.cabal, or download it on demand or something. We need a single buck2 binary installed somewhere - ghcup could manage that when installing cabal, perhaps.

I have two projects that I would love to be able to build this way: hsthrift and Glean. And this opens up a real possibility that GHC could be just a Cabal project. (yes we can extend Cabal enough to build GHC today, but IMO that’s not the way to go.)

There’s a bit of an issue with projects vs. packages. Making BUCKfiles that work both in a project setting and a single package build needs to be worked out (maybe a package == a buck2 “cell”, I’m not sure).

Design 2: Buck2 is an optional build tool. All build customisation must be done with Cabal hooks. The advantage of this route is that in principle we have a choice of build tools - Ninja, Buck2, Bazel, Cabal itself. But TBH I think this is a non-starter, nobody really wants to use Cabal hooks to write build system code (do they?).

1 Like
1 Like

Is there a particular reason it should be many binaries specifically, rather than one binary with many separate modes (not unlike how it is today – git is another example (with even more modes)) implemented on top of a well-factored API. In order to split cabal into many binaries it would have to have a well-factored API. Once that’s been achieved, what’s the additional benefit of many binaries?

It doesn’t really matter much. There’s two parts of the story:

  1. a “shell API”, which means I could build my own cabal-install trivially by invoking sub-parts and only change one part (e.g. you could swap the installation part with “make a tarball out of this”… something we ran into in stable-haskell and it’s surprisingly hard with current cabal)… everything is just binary1 | binary2 | binary3
  2. forcing boundaries. Having seperate executables (living in separate packages) just seems a more radical approach to me. But you could achieve that in all sorts of different ways and could then provide the shell API via subcommands, symlink based dispatching, whatever
1 Like

I think the most promising direction is a unified CLI with a stable, versioned build-plan API and pluggable executors.

Users should still get a simple cabal build experience, while tools like Buck2, Bazel, or a native executor can consume the same explicit build graph underneath.

The key abstraction should be “describe what needs to be built,” not “run arbitrary build-time IO.” That preserves Cabal’s package semantics while making caching, remote execution, and integration with larger multi-language build systems much easier.

1 Like