Generate Unicode Tables from the UCD
Status
Accepted
Context
Ferroni originally generated its Unicode tables from Oniguruma's generated C data. Oniguruma's final source snapshot contains Unicode 16.0 data. Continuing to use that snapshot would leave Ferroni unable to adopt newer Unicode releases, even though the Unicode Character Database (UCD) continues to be published.
The project still follows the structural parity goal in ADR-001. Refreshing the Unicode inputs does not change the regex engine's control flow; it updates the versioned data used by the existing property, case-fold, grapheme-break, and word-break code.
Decision
Generate the Rust tables directly from the official, versioned UCD archives. The source URLs and SHA-256 checksums are recorded in unicode_data.toml. The preparation script rejects archives whose checksum does not match the pin. The generator verifies that the UCD and emoji data versions agree before writing the four tables in src/unicode.
The baseline migration was checked against Unicode 16.0. The generated property, case-fold, grapheme-break, and word-break data match the previously checked-in tables exactly. The generated provenance comments and the Extended_Pictographic ctype index declaration differ; the index is now emitted from the same property ordering as the tables instead of being hard-coded in the break algorithm. scripts/check_unicode_16.py preserves the table-body comparison with committed SHA-256 values and checks the old index value.
Ferroni now uses Unicode 17.0.0 data. Its Unicode property lookup recognizes 902 names, up from 886. This is an intentional data-level divergence from the frozen Oniguruma snapshot. The regex behavior and table layout continue to follow the existing port.
The generator sorts each property's ranges before merging them. The
DerivedCoreProperties.txt file lists Indic_Conjunct_Break (InCB) in three
sections (Linker, Consonant, Extend) that are not in code point order. Merging
them unsorted drops ranges: with Unicode 17.0, \p{InCB} covered only 720 of
its 3148 code points. Oniguruma's own generator has the same defect, and its
Unicode 16.0 table keeps 293 of 398 ranges (1717 of 2438 code points).
Ferroni's complete InCB table is an intentional divergence from C
Oniguruma. No other property changes, because every other property is listed
in code point order. The Unicode 16.0 baseline hash for the property table was
updated accordingly; the case-fold, grapheme-break, and word-break baselines
are unchanged.
Consequences
- Updating Unicode data no longer depends on an Oniguruma checkout.
- The UCD archive is a reproducible input: its version, source URL, and checksum are all pinned.
- A Unicode release update regenerates the property, case-fold, grapheme-break, and word-break tables together.
- Table generation requires Python 3.11 or later and rustfmt.
- Existing behavior for Unicode 16.0 data remains guarded by the baseline verifier.
\p{InCB}matches every Indic_Conjunct_Break code point, which C Oniguruma does not; patterns relying on the C table's gaps behave differently.