# \[Blog\] 5x speedup changing break to elemIndex!

**URL:** <https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946>\
**Category:** Links\
**Created:** [April 18, 2026, 5:32am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946 "2026-04-18T05:32:23Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![brandonchinn178](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/brandonchinn178/32/5103_2.png) [@brandonchinn178](https://discourse.haskell.org/u/brandonchinn178)\
**Post date:** [April 18, 2026, 5:32am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/1 "2026-04-18T05:32:23Z")

</div>

> **[Optimizing xreferee with elemIndex](https://brandonchinn178.github.io/posts/2026/04/17/optimizing-xreferee-with-elemindex/)**
>
> xreferee is a linter that checks that every @(ref:foo) reference in a git repository corresponds to a #(ref:foo) anchor somewhere in the repository. It delegates most of the search to git grep, but there's some parsing logic to parse git grep's...

---

<div class="post-metadata">

**Author:** ![pmidden](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/pmidden/32/2426_2.png) [@pmidden](https://discourse.haskell.org/u/pmidden)\
**Post date:** [April 19, 2026, 7:39am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/2 "2026-04-19T07:39:40Z")

</div>

Would be interesting to know how something like attoparsec fares here as a little higher-level abstraction.

---

<div class="post-metadata">

**Author:** ![prophet](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/prophet/32/5031_2.png) [@prophet](https://discourse.haskell.org/u/prophet)\
**Post date:** [April 19, 2026, 10:34am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/3 "2026-04-19T10:34:19Z")

</div>

Very cool! This is a nice general optimization.

In your specific case, I think there is a much easier solution though. `git grep` has already parsed the exact match you’re looking for and you only need to re-parse every line because your git grep returns the _entire_ line.  
If you tell it to return only the exact matches (with `git grep -o` and a pattern that includes the contents of the references like with -E: `(@|#)\(ref:[^)]*\)`, then you will get one result per match (instead of one per line) and only need to trimm off the first and last character

---

<div class="post-metadata">

**Author:** ![brandonchinn178](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/brandonchinn178/32/5103_2.png) [@brandonchinn178](https://discourse.haskell.org/u/brandonchinn178)\
**Post date:** [April 19, 2026, 3:28pm UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/4 "2026-04-19T15:28:04Z")

</div>

Thanks for the reminder! An old version of the original algorithm used -o, but we removed it because the git that ships with Centos 7 was too old and didn’t have the flag. But I am using `--columns` which is newer than `-o`, so let me add `-o` and document the minimum git version required. Thanks!

---

<div class="post-metadata">

**Author:** ![unhammer](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/unhammer/32/580_2.png) [@unhammer](https://discourse.haskell.org/u/unhammer)\
**Post date:** [April 20, 2026, 7:28am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/5 "2026-04-20T07:28:07Z")

</div>

Would it make sense for the docs for `break` to have a note on performance, or mention that there are much faster ways to break on single specific characters?

---

<div class="post-metadata">

**Author:** ![Kleidukos](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/kleidukos/32/1213_2.png) [@Kleidukos](https://discourse.haskell.org/u/Kleidukos)\
**Post date:** [April 20, 2026, 8:07am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/6 "2026-04-20T08:07:41Z")

</div>

Yes it would make sense

---

<div class="post-metadata">

**Author:** ![Bodigrim](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/bodigrim/32/1457_2.png) [@Bodigrim](https://discourse.haskell.org/u/Bodigrim)\
**Post date:** [April 20, 2026, 6:39pm UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/7 "2026-04-20T18:39:47Z")

</div>

I think the migration from `Text` to `ByteString` could provide only marginal gains and probably was not worth it. UTF-8 decoding is implemented with SIMD instructions, so it is very fast, especially on the happy path. And one can use `decodeLenient` to skip decoding failures.

You could have looked inside `Text` to find `ByteArray` suitable for [`Data.Text.Internal.ArrayUtils.memchr`](https://hackage-content.haskell.org/package/text-2.1.4/docs/Data-Text-Internal-ArrayUtils.html#v:memchr) (yes, suspiciously enough `text` already uses `memchr`). For a more ergonomic solution I’d welcome PRs for [Add RULE from break to breakOn · Issue #695 · haskell/text · GitHub](https://github.com/haskell/text/issues/695) and [Search of a singleton needle should use memchr · Issue #696 · haskell/text · GitHub](https://github.com/haskell/text/issues/696).

---

<div class="post-metadata">

**Author:** ![brandonchinn178](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/brandonchinn178/32/5103_2.png) [@brandonchinn178](https://discourse.haskell.org/u/brandonchinn178)\
**Post date:** [April 21, 2026, 6:29am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/8 "2026-04-21T06:29:11Z")

</div>

Well the migration from `Text` to `ByteString` unblocked using `elemIndex`, so it’s worth it now. In a future where `Text` provides an equivalent of `elemIndex`, sure. No, I’d rather not use that internal `memchr` function 🙂

---

<div class="post-metadata">

**Author:** ![brandonchinn178](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/brandonchinn178/32/5103_2.png) [@brandonchinn178](https://discourse.haskell.org/u/brandonchinn178)\
**Post date:** [April 21, 2026, 6:29am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/9 "2026-04-21T06:29:56Z")

</div>

So fun story: I tried this, and found a git bug 😄

> **[git grep bug with --column and --only-matching](https://lore.kernel.org/git/CAGANf=dkRgFp+bEkB5f8QBeiR3m+3WE8sKqT9vKstkGHqbxA3A@mail.gmail.com/T/#u)**

---

<div class="post-metadata">

**Author:** ![Bodigrim](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/bodigrim/32/1457_2.png) [@Bodigrim](https://discourse.haskell.org/u/Bodigrim)\
**Post date:** [August 12, 2026, 9:40pm UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/10 "2026-08-12T21:40:03Z")

</div>

> [@Bodigrim](#):
>
> For a more ergonomic solution I’d welcome PRs for [Add RULE from break to breakOn · Issue #695 · haskell/text · GitHub](https://github.com/haskell/text/issues/695) and [Search of a singleton needle should use memchr · Issue #696 · haskell/text · GitHub](https://github.com/haskell/text/issues/696) .

…and thanks to @Lysxia both optimizations have been implemented now.

---

<div class="post-metadata">

**Author:** ![Ambrose](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/ambrose/32/5672_2.png) [@Ambrose](https://discourse.haskell.org/u/Ambrose)\
**Post date:** [August 13, 2026, 7:24am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/11 "2026-08-13T07:24:40Z")

</div>

tfw you gotta use `RULES`

---

<div class="post-metadata">

**Author:** ![enobayram](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/enobayram/32/1815_2.png) [@enobayram](https://discourse.haskell.org/u/enobayram)\
**Post date:** [August 13, 2026, 8:54am UTC](https://discourse.haskell.org/t/blog-5x-speedup-changing-break-to-elemindex/13946/12 "2026-08-13T08:54:11Z")

</div>

Haskell: The language where bad boys play by the `RULES`
