# \[ANN\] unicode-data-0.3.0: APIs to efficiently access the Unicode character database

**URL:** <https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861>\
**Category:** Uncategorized\
**Created:** [January 2, 2022, 11:25am UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861 "2022-01-02T11:25:43Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![Wismill](https://avatars.discourse-cdn.com/v4/letter/w/ccd318/32.png) [@Wismill](https://discourse.haskell.org/u/Wismill)\
**Post date:** [January 2, 2022, 11:25am UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/1 "2022-01-02T11:25:43Z")

</div>

On behalf of the maintainers team I’m happy to announce [`unicode-data-0.3.0`](https://hackage.haskell.org/package/unicode-data-0.3.0).

`unicode-data` provides Haskell APIs to efficiently access the latest **Unicode character database**.  
It is up to [_5 times faster than `base:Data.Char`_](https://github.com/composewell/unicode-data#performance).

This release features:

- Support for big-endian architectures.
- Support for `General_Category` property.
- All the predicates and case mapping functions found in `base:Data.Char`.

It follows the 0.2.0 release which added support for:

- [Unicode 14.0.0.](https://www.unicode.org/versions/Unicode14.0.0/)
- Unicode Identifier and Pattern Syntax.

See the [complete ChangeLog](https://hackage.haskell.org/package/unicode-data-0.3.0/changelog) for details.

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [January 2, 2022, 6:19pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/2 "2022-01-02T18:19:20Z")

</div>

> [@Wismill](#):
>
> It is up to [_5 times faster than `base:Data.Char`_](https://github.com/composewell/unicode-data#performance).

Is there a reason why we can’t make `Data.Char` just as fast?

---

<div class="post-metadata">

**Author:** ![Wismill](https://avatars.discourse-cdn.com/v4/letter/w/ccd318/32.png) [@Wismill](https://discourse.haskell.org/u/Wismill)\
**Post date:** [January 2, 2022, 6:36pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/3 "2022-01-02T18:36:36Z")

</div>

We use a different approach: `Data.Char` relies on FFI (see `GHC.Unicode`) while `unicode-data` uses only _pure_ Haskell (bitmaps and simple functions). Maybe @Bodigrim or @harendra could explain this better?

---

<div class="post-metadata">

**Author:** ![chreekat](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/chreekat/32/2669_2.png) [@chreekat](https://discourse.haskell.org/u/chreekat)\
**Post date:** [January 2, 2022, 6:45pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/4 "2022-01-02T18:45:18Z")

</div>

Maybe I’m dense, but that sounds like a “how” rather than a “why”. Put it another way, is there a reason Data.Char can’t be reimplemented to use your approach? A pure-Haskell approach sounds better, anyway. 🙂

Anyway, good work, thanks for the library!

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [January 2, 2022, 6:49pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/5 "2022-01-02T18:49:53Z")

</div>

Indeed, pure Haskell _and_ faster sounds like a clear win for everyone (especially people porting Haskell to odd environments). Sounds like folding this into `base`, or reversing the dependency direction, is worth discussing!

How does code size change?

---

<div class="post-metadata">

**Author:** ![Bodigrim](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/bodigrim/32/1457_2.png) [@Bodigrim](https://discourse.haskell.org/u/Bodigrim)\
**Post date:** [January 2, 2022, 6:59pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/6 "2022-01-02T18:59:57Z")

</div>

It is technically possible to fold `unicode-data` into `base`, but I would advise against. The reason is that Unicode is an evolving standard, and it is desirable to have an ability to update to the latest version without upgrading your compiler. A standalone package provides such possibility, but `base` does not - you are bound to whichever Unicode version was wired in. I suspect not all developers fully realise that the behaviour of Unicode-aware application depends on GHC used to build it. In fact, it would be better to deprecate Unicode API from `Data.Char` and refer users to `unicode-data` or similar packages.

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [January 2, 2022, 7:23pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/7 "2022-01-02T19:23:03Z")

</div>

Well, `Data.Char` is currently there, and is widely used. As long as it is there and not deprecated, surely changing an FFI implementation to a faster and pure implementation seems desirable – the problems with outdated unicode standards (which probably not all developers care about) are orthogonal to that. And those developers who do care can of course still use such a dedicated library.

---

<div class="post-metadata">

**Author:** ![ketzacoatl](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/ketzacoatl/32/1203_2.png) [@ketzacoatl](https://discourse.haskell.org/u/ketzacoatl)\
**Post date:** [January 3, 2022, 5:54am UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/8 "2022-01-03T05:54:21Z")

</div>

IMHO, we should not restrict the correctness or capabilities of a basic library such as `unicode-data` for the GHC/base versioning issues. We should instead separate the GHC-dependent and independent stuff in `base`, and otherwise de-couple the versioning there. I know this is more complicated, but I’m sure we can figure it out (for example, the core parts of unicode-data that need to be in base can be, with “faster moving stuff” in a separate package). But even so, let’s clean up GHC internals so we stop crippling ourselves.

---

<div class="post-metadata">

**Author:** ![Bodigrim](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/bodigrim/32/1457_2.png) [@Bodigrim](https://discourse.haskell.org/u/Bodigrim)\
**Post date:** [January 3, 2022, 5:25pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/9 "2022-01-03T17:25:39Z")

</div>

@nomeata there is more motivation to use a proper solution, when it is both faster and correct than when it is just correct 😉 But anyway `unicode-data` is open-source, so nothing prevents a motivated individual to put an effort and merge it into `base`, I believe.

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [January 3, 2022, 6:07pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/10 "2022-01-03T18:07:54Z")

</div>

> [@Bodigrim](#):
>
> nothing prevents a motivated individual to put an effort and merge it into `base`, I believe.

Is that an official approval by the CLC? 😉  
Or who’d be the one to tell such an motivated individual that such a change would be welcome?

@Wismill, do you know off hand if your approach can replace all of the following (from `include/WCsubst.h`)

```haskell
HsInt u_iswupper(HsInt wc);
HsInt u_iswdigit(HsInt wc);
HsInt u_iswalpha(HsInt wc);
HsInt u_iswcntrl(HsInt wc);
HsInt u_iswspace(HsInt wc);
HsInt u_iswprint(HsInt wc);
HsInt u_iswlower(HsInt wc);

HsInt u_iswalnum(HsInt wc);

HsInt u_towlower(HsInt wc);
HsInt u_towupper(HsInt wc);
HsInt u_towtitle(HsInt wc);

HsInt u_gencat(HsInt wc);

```

---

<div class="post-metadata">

**Author:** ![Bodigrim](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/bodigrim/32/1457_2.png) [@Bodigrim](https://discourse.haskell.org/u/Bodigrim)\
**Post date:** [January 3, 2022, 9:04pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/11 "2022-01-03T21:04:23Z")

</div>

> Or who’d be the one to tell such an motivated individual that such a change would be welcome?

It depends on how and where it is impemented. E. g., you can just rewrite [Files · master · Glasgow Haskell Compiler / GHC · GitLab](https://gitlab.haskell.org/ghc/ghc/-/blob/master/libraries/base/cbits/WCsubst.c) using lookup tables (similar to `unicode-data`) instead of binary search. This will give you the ultimate performance, ~~and no need for CLC approval, because this is not even Haskell~~ (this C module is a part of `base` and technically still falls under CLC supervision, but I do not expect a lot of fight over an autogenerated file).

---

<div class="post-metadata">

**Author:** ![harendra](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/harendra/32/1143_2.png) [@harendra](https://discourse.haskell.org/u/harendra)\
**Post date:** [January 11, 2022, 5:36am UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/12 "2022-01-11T05:36:06Z")

</div>

> [@nomeata](#):
>
> do you know off hand if your approach can replace all of the following (from `include/WCsubst.h`)

Yes, all of these are implemented. In fact, unicode-data is a drop-in replacement for `Data.Char` except for a couple of Show/Read related functions.

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [February 22, 2022, 7:11pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/13 "2022-02-22T19:11:43Z")

</div>

On [javascript backend · Wiki · Glasgow Haskell Compiler / GHC · GitLab](https://gitlab.haskell.org/ghc/ghc/-/wikis/javascript-backend) I spotted

> What is the plan for handling the various C bits in base/text/bytestring/etc.? Specifically, has there been any motion on discussions with upstreams regarding upstreaming the Javascript implementations?
> 
> We plan to propose patches for these libraries. We can use CPP to condition JS specific code to `HOST_OS=ghcjs` and similarly in `.cabal` files.

I guess a pure Haskell solution for `Data.Char` would help with that kind of work.

---

<div class="post-metadata">

**Author:** ![nomeata](https://sea2.discourse-cdn.com/flex002/user_avatar/discourse.haskell.org/nomeata/32/1188_2.png) [@nomeata](https://discourse.haskell.org/u/nomeata)\
**Post date:** [April 11, 2022, 1:00pm UTC](https://discourse.haskell.org/t/ann-unicode-data-0-3-0-apis-to-efficiently-access-the-unicode-character-database/3861/14 "2022-04-11T13:00:02Z")

</div>

I started a discussion about how `base` can benefit here: [https://gitlab.haskell.org/ghc/ghc/-/issues/21375](https://gitlab.haskell.org/ghc/ghc/-/issues/21375). @wismill, your thoughts will be appreciated there
