PR 173: Full Oniguruma Comparison
This report evaluates general regex tasks and scanner workloads for
PR #173, the opt-in
match-cache reconstruction. The initial reference run covers 72 timings across
29 paired Ferroni/C scenarios, with the regex crate in 14 compatible cases.
A follow-up adds seven everyday regex tasks measured with all three engines.
Separate runs measure the opt-in cache on three documents and compare process
memory. Syntax highlighting is one integration workload; it does not stand in
for the general engine.
The general-regex follow-up and initial reference run both compile the new failed-state match cache out. Its benefit on pathological inputs and review requirements are documented in the match-cache evaluation. The separate cache-on/off document table measures the cost of opting in on ordinary documents. It does not attribute existing Ferroni advantages to the new cache.
What the comparison covers
The primary question is how Ferroni performs against the C engine whose
behavior it ports: Oniguruma. The main tables compare those two engines,
including patterns that use Oniguruma features unavailable in regex.
| Capability exercised here | Ferroni | Oniguruma | Rust regex |
|---|---|---|---|
| Literals, classes, quantifiers, alternation | Supported | Supported | Supported |
| Capture groups and the tested Unicode patterns | Supported | Supported | Supported |
| Lookahead and lookbehind | Supported and measured | Supported and measured | Not supported |
| Backreferences | Supported and measured | Supported and measured | Not supported |
| RegSet/scanner workload in this harness | Measured | Measured | No equivalent path measured |
This is a scope map for the benchmarks, not a complete compatibility matrix.
regex deliberately excludes lookarounds and backreferences.
The seven everyday tasks below all use shared syntax; they do not exercise
those differences. A result on that subset cannot establish interchangeability
for an existing Oniguruma pattern set.
The shared-syntax appendix retains every
regex measurement as context for applications whose patterns fit that subset.
Unsupported syntax and an unmeasured API path are separate limitations; neither
is a timing result. There is no winner count across engines with different
feature coverage.
Everyday regex tasks: Ferroni vs Oniguruma
This follow-up focuses on the general engine: three validation batches and four complete text-processing tasks, with no TextMate scanner. Ferroni has the lower point estimate than C in all seven cases, by 1.08–1.98x. The scanner speedups are not a useful estimate for ordinary regex work.
| Task (whole batch/text) | Ferroni | Oniguruma | C / Ferroni |
|---|---|---|---|
| Email shape, 64 inputs | 6.722 us | 12.843 us | 1.91x |
| UUID shape, 64 inputs | 11.079 us | 15.773 us | 1.42x |
| Number syntax, 64 inputs | 8.034 us | 14.401 us | 1.79x |
| Access log, 64 records / 6 fields | 32.410 us | 34.866 us | 1.08x |
| Extract 64 URLs | 9.871 us | 19.546 us | 1.98x |
| Extract 384 Unicode words | 63.527 us | 81.591 us | 1.28x |
| Redact 128 email addresses | 78.465 us | 87.898 us | 1.12x |
Times cover the whole batch or text, not one match. Bold marks the lowest point estimate between Ferroni and C, not statistical significance. The validators repeat eight labeled inputs eight times: 32 accepted and 32 rejected, including early rejections and failures near the end. They measure boolean results without capture output. Email and UUID are simplified shape checks, not complete RFC or UUID-version validation; the number case checks signed decimal/exponent syntax without numeric conversion.
The text tasks use deterministic synthetic inputs from
general_regex.rs:
- Access log: 5,964 bytes, 64 matching records, six captured fields each, plus 16 nonmatching messages.
- URL extraction: 5,196 bytes and 64 HTTP/HTTPS links amid plain text.
- Unicode words: 3,648 bytes and 384 words using Latin, Greek, Cyrillic, Arabic,
Han characters, and a decomposed accent. The pattern is
\p{L}[\p{L}\p{M}]*. - Email redaction: 5,218 bytes, 128 addresses, and similar-looking nonmatches.
Every match becomes
[email]through the same output builder for all engines.
Each text operation collects every match and capture into byte-offset vectors. Capture scratch and result allocation are inside timing; the region or capture-location storage is reused between searches within one operation. The redaction row includes the common output builder and is not a comparison of native replacement APIs. Compilation, input generation, and correctness checks are outside timing. Patterns are compiled once and reused. These inputs contain no empty matches; empty-match iteration policy is outside this scope.
Before timing, all three engines agree on each labeled validation result and
every extracted match/capture boundary. The preflight also checks expected
match counts, concrete log-field values, and the exact redacted output.
The benchmark smoke run covers all 96 cases with ffi,match-cache enabled.
Access-log capture and URL extraction were repeated separately with the same configuration. Ferroni remained faster than C on these two inputs. The appendix retains the third engine's repeat timings as shared-syntax context.
| Repeat (whole text) | Ferroni | Oniguruma | C / Ferroni |
|---|---|---|---|
| Access log | 31.591 us | 33.240 us | 1.05x |
| URL extraction | 10.080 us | 19.530 us | 1.94x |
The main table retains the original run. These are a few representative tasks, not an application-traffic distribution or a cross-platform engine ranking. Backreferences, lookarounds, single-pattern compilation, and long no-match searches remain in the initial reference tables below; syntax highlighting remains a separate integration workload.
The follow-up was measured on 2026-09-26 at harness commit
03a6ba69c5be294e77ca20ce9f0a08db4973f0a4, with the same machine, compilers,
dependencies, input pins, and 30-sample / 500 ms warmup / 4 s measurement
settings listed below. Engine sources are unchanged from 9f18790.
It uses --features ffi, so the opt-in match cache is compiled out. Each case
runs Ferroni, then C, then regex; the repeats retain that fixed order.
This is a separate run from the original scanner/feature suite.
The follow-up raw artifact contains all 21 initial measurements and six repeat measurements, 95% confidence intervals, individual samples, exact patterns, workload sizes, and source hashes.
cargo bench --locked --features ffi --bench battle_bench -- \
general_regex --noplot --save-baseline pr173-general-regex
# Render before repeats overwrite Criterion's `new` directories.
python3 scripts/gen_battle_tables.py --general-only
cargo bench --locked --features ffi --bench battle_bench -- \
access_log_captures --noplot --save-baseline pr173-general-log-repeat
cargo bench --locked --features ffi --bench battle_bench -- \
url_extraction --noplot --save-baseline pr173-general-url-repeatInitial reference-run measurement context
| Field | Value |
|---|---|
| Date | 2026-09-26 |
| Ferroni/harness commit | 719a51f9e0caba622e90470f493f5f5b89848d64 |
| Engine source commit | 9f1879007d735e791fc33c69133920ff22fc1640 (unchanged by the harness update) |
| Host | Mac13,2, Apple M1 Ultra, 64 GiB RAM |
| OS / target | macOS 27.0 (26A428), aarch64-apple-darwin |
| Rust | rustc 1.96.0 (ac68faa20), LLVM 22.1.2 |
| C compiler | Apple clang 21.0.0 (clang-2100.3.34.2) |
| Profiles | Optimized Rust bench profile, thin LTO; C Oniguruma at -O3 |
| Dependencies | Criterion 0.8.2; regex 1.13.1; checked-in Cargo.lock |
| Criterion | 30 samples, 500 ms warmup, 4 s measurement per case |
| Oniguruma | f95747b462de672b6f8dbdeb478245ddf061ca53 |
| Full grammars | TypeScript: 279 patterns; CSS: 117; Rust: 81 |
Criterion extends slow cases when four seconds would not collect 30 samples.
Tables report the slope estimate when available, otherwise the mean. The
raw artifact
contains all 83 timing measurements (72 reference, two CSS repeat, nine cache
comparison), 95% confidence intervals, individual sample durations/iteration
counts, all three memory runs, source hashes, and the input pins from
battle_inputs.toml.
The full run, repeat, cache comparison, and memory processes ran sequentially. This is a developer-workstation measurement with a fixed engine order, not a cross-platform or randomized campaign. Small percentage differences require repetition before becoming release criteria.
Equivalent work and correctness checks
The harness previously requested capture output from C in timed text and single-pattern searches while Ferroni received no output region. This run requests no capture output from either engine. Separate, untimed assertions compare exact match positions and raw capture bounds. Scanner timings include captures for both engines; their preflight compares pattern indices and every capture boundary along the complete fixture trace. These checks run in optimized builds too, including the cache-enabled lane.
The scanner fixtures are ASCII, so their byte and UTF-16 offsets coincide.
This validates these benchmark inputs; it does not replace the broader Unicode,
differential, and fuzz evidence in the cache evaluation. regex is syntax-limited
context, and its find API returns both start and end positions.
The older published benchmark tables remain a dated snapshot on a different machine and toolchain, with the previous harness. Do not calculate a version-to-version regression from the two pages.
Initial reference suite
In the initial run, default Ferroni has a lower point estimate than C Oniguruma in all 23 search/scanner scenarios and in three of six compilation scenarios. Whole-document tokenization is approximately 3.0x (TypeScript), 31.7x (CSS), and 10.3x (Rust) faster in the full run. CSS has a noisy initial sample; the targeted repeat and both sets of raw samples are retained below. These ratios describe this suite, not a universal speedup.
Oniguruma still compiles the simple literal, lookbehind, and full Rust grammar
faster. Shared-syntax timings for regex are retained in the appendix rather
than treated as a ranking of engines with equivalent capabilities.
Bold marks the lower Ferroni/C point estimate, not statistical significance. Factors are C time divided by Ferroni time; values above one favor Ferroni.
Text search and log scanning
| Scenario | Ferroni | Oniguruma | C / Ferroni |
|---|---|---|---|
| Literal in 50 KB | 70.279 ns | 162.110 ns | 2.3x |
| No match, 50 KB | 1.554 us | 9.434 us | 6.1x |
| No match, 10 KB | 361.350 ns | 1.881 us | 5.2x |
| Field extract, 50 KB | 94.270 ns | 172.352 ns | 1.83x |
| Timestamp, 50 KB | 95.618 ns | 192.295 ns | 2.0x |
| RegSet multi-pattern (5) | 100.562 ns | 399.525 ns | 4.0x |
Pattern matching
| Category | Ferroni | Oniguruma | C / Ferroni |
|---|---|---|---|
| Literal exact | 102.612 ns | 157.929 ns | 1.54x |
| Quantifier greedy | 62.308 ns | 200.063 ns | 3.2x |
| Lookaround combined | 87.678 ns | 248.714 ns | 2.8x |
Unicode \p{Greek}+ | 103.399 ns | 213.873 ns | 2.1x |
Backref (\w+) \1 | 84.129 ns | 174.816 ns | 2.1x |
| Case-insensitive phrase | 104.720 ns | 185.957 ns | 1.78x |
| Alternation, 2 branches | 66.211 ns | 160.676 ns | 2.4x |
| Alternation, 10 branches | 55.980 ns | 242.930 ns | 4.3x |
| Named capture date | 221.172 ns | 260.521 ns | 1.18x |
Compilation
| Pattern | Ferroni | Oniguruma | C / Ferroni |
|---|---|---|---|
| Literal | 507.053 ns | 493.223 ns | 0.97x |
| Named capture | 4.113 us | 6.337 us | 1.54x |
| Lookbehind | 1.161 us | 611.973 ns | 0.53x |
Scanner with full Shiki TextMate grammars
| Scenario | Ferroni | Oniguruma | Factor |
|---|---|---|---|
| TypeScript (279 patterns) | |||
| Compile | 11.088 ms | 17.454 ms | 1.57x |
| First match, short line | 85.084 ns | 25.569 us | 300.5x |
| Tokenize full line | 2.707 us | 204.592 us | 75.6x |
| CSS (117 patterns) | |||
| Compile | 14.870 ms | 18.541 ms | 1.25x |
| Tokenize (multi-line) | 221.197 us | 14.712 ms | 66.5x |
| Rust (81 patterns) | |||
| Compile | 319.448 us | 195.092 us | 0.61x |
| First match | 172.934 ns | 5.505 us | 31.8x |
| Tokenize full line | 4.930 us | 80.209 us | 16.3x |
Scanner on whole documents, line by line
| Document | Ferroni | Oniguruma | Factor |
|---|---|---|---|
| TypeScript (279 patterns), 28 lines | 1.035 ms | 3.129 ms | 3.0x |
| CSS (117 patterns), 19 lines | 100.494 us | 3.190 ms | 31.7x |
| Rust (81 patterns), 31 lines | 105.659 us | 1.086 ms | 10.3x |
The scanner-highlighting rows reuse the same OnigString. Ferroni's existing
per-string memo avoids repeated fallback work, while the C scanner wrapper
consults its per-pattern cache only at 1,000 bytes or more. All those fixtures
are shorter. The large warm-path ratios therefore describe these two wrapper
policies as well as the engines.
The whole-document rows tokenize each fixture line by line, preserving trailing newlines and changing string identity per line. Compilation is excluded, and the per-line strings are prepared before timing. These are scanner workloads over full grammar pattern sets, not end-to-end Shiki/TextMate rendering with grammar-state transitions.
The initial CSS Ferroni estimate was 100.494 us, with a 95% confidence interval of 90.240–115.971 us. A targeted repeat with the same settings measured 89.059 us for Ferroni and 3.126 ms for C (35.1x), with a much narrower Ferroni interval of 88.933–89.158 us. The main table deliberately retains the complete run instead of substituting the faster repeat. No single aggregate combines cold, warm, search, and compilation ratios.
Opt-in cache on ordinary documents
This separate ffi,match-cache build measures runtime cache off, runtime cache
on (default 16 MiB budget and adaptive activation), and C in the same executable.
The inputs, trace validation, and timing settings match the document rows above.
Opt-in cost is (cache on / cache off - 1) within this feature-enabled build;
do not calculate that cost across different executables.
| Document | Cache off | Cache on | Oniguruma | Opt-in cost |
|---|---|---|---|---|
| TypeScript | 1.126 ms | 1.141 ms | 3.162 ms | +1.38% |
| CSS | 99.183 us | 102.197 us | 3.288 ms | +3.04% |
| Rust | 120.742 us | 128.018 us | 1.113 ms | +6.03% |
The cache protects eligible expensive failures; ordinary grammar documents still pay bookkeeping costs. This run does not justify default enablement. The feature-disabled release comparison and pathological cases remain in the separate evaluation.
Process memory
The memory harness compiles the same full TypeScript grammar and scans a
deterministic 1,204,992-byte source containing 16,002 lines. It requests
the first match from each line, not complete tokenization. Both engines reported
16,002 scanned and matched lines in each of three runs. Each engine runs in its
own process; Ferroni uses the default build without match-cache.
| Engine | Phase | Median peak RSS | Range (three runs) |
|---|---|---|---|
| Ferroni | Compile | 15,106,048 B | 15,089,664–15,122,432 B |
| Ferroni | Compile + line scan | 15,204,352 B | 15,187,968–15,220,736 B |
| Oniguruma | Compile | 15,728,640 B | 15,482,880–15,990,784 B |
| Oniguruma | Compile + line scan | 15,876,096 B | 15,679,488–16,187,392 B |
These are process high-water RSS values from getrusage, including the driver,
input, grammar storage, allocator, and compiled engine. The scan value includes
the preceding compile phase. They are not engine heap sizes, incremental scan
allocations, or measurements of the opt-in cache's 16 MiB budget. See the
memory methodology for the harness boundaries.
Shared-syntax context only: regex
These measurements answer a narrower question: how quickly each engine runs
the particular patterns supported by all three. They do not compare the full
Oniguruma feature set or establish whether regex can replace Ferroni.
regex excludes lookarounds and backreferences
as part of its design. Its timings remain useful when an application's
patterns fit that subset. Capture groups and Unicode are supported; those
features alone do not require an Oniguruma-compatible engine.
The original measurements are retained below without winner counts or an overall ranking. Only measured shared-syntax cases appear here. Lookarounds, backreferences, and the unmeasured RegSet/scanner path are excluded explicitly, not assigned a slow or zero result. Times and measurement conditions are unchanged from their corresponding main tables.
Shared syntax: everyday regex tasks
| Task (whole batch/text) | Ferroni | Oniguruma | regex (shared syntax only) |
|---|---|---|---|
| Email shape, 64 inputs | 6.722 us | 12.843 us | 1.522 us |
| UUID shape, 64 inputs | 11.079 us | 15.773 us | 1.769 us |
| Number syntax, 64 inputs | 8.034 us | 14.401 us | 1.033 us |
| Access log, 64 records / 6 fields | 32.410 us | 34.866 us | 35.726 us |
| Extract 64 URLs | 9.871 us | 19.546 us | 13.364 us |
| Extract 384 Unicode words | 63.527 us | 81.591 us | 32.798 us |
| Redact 128 email addresses | 78.465 us | 87.898 us | 13.915 us |
Shared syntax: everyday task repeats
| Repeat (whole text) | Ferroni | Oniguruma | regex (shared syntax only) |
|---|---|---|---|
| Access log | 31.591 us | 33.240 us | 35.684 us |
| URL extraction | 10.080 us | 19.530 us | 13.377 us |
Shared syntax: text search and log scanning
| Scenario | Ferroni | Oniguruma | regex (shared syntax only) |
|---|---|---|---|
| Literal in 50 KB | 70.279 ns | 162.110 ns | 9.757 ns |
| No match, 50 KB | 1.554 us | 9.434 us | 1.459 us |
| No match, 10 KB | 361.350 ns | 1.881 us | 296.226 ns |
| Field extract, 50 KB | 94.270 ns | 172.352 ns | 54.945 ns |
| Timestamp, 50 KB | 95.618 ns | 192.295 ns | 63.992 ns |
Shared syntax: pattern matching
| Category | Ferroni | Oniguruma | regex (shared syntax only) |
|---|---|---|---|
| Literal exact | 102.612 ns | 157.929 ns | 11.091 ns |
| Quantifier greedy | 62.308 ns | 200.063 ns | 64.364 ns |
Unicode \p{Greek}+ | 103.399 ns | 213.873 ns | 58.792 ns |
| Case-insensitive phrase | 104.720 ns | 185.957 ns | 60.377 ns |
| Alternation, 2 branches | 66.211 ns | 160.676 ns | 45.972 ns |
| Alternation, 10 branches | 55.980 ns | 242.930 ns | 20.073 ns |
| Named capture date | 221.172 ns | 260.521 ns | 44.633 ns |
Shared syntax: compilation
| Pattern | Ferroni | Oniguruma | regex (shared syntax only) |
|---|---|---|---|
| Literal | 507.053 ns | 493.223 ns | 2.896 us |
| Named capture | 4.113 us | 6.337 us | 208.787 us |
Reproducing this snapshot
To reproduce the original 72-measurement suite exactly, use harness commit
719a51f. The current full suite also includes general_regex; the follow-up
commands above isolate that addition. Run each timing command to completion
before starting the next one. Group configuration in battle_bench.rs fixes
the Criterion settings listed above.
./scripts/prepare-oniguruma-sources.sh
cargo bench --locked --features ffi --bench battle_bench -- \
--noplot --save-baseline pr173-reference
# Render before the filtered runs overwrite Criterion's `new` directories.
python3 scripts/gen_battle_tables.py
cargo bench --locked --features ffi --bench battle_bench -- \
scanner_documents/css_117 --noplot --save-baseline pr173-css-repeat
cargo bench --locked --features ffi,match-cache --bench battle_bench -- \
scanner_documents --noplot --save-baseline pr173-cache-documents
# Three sequential invocations, one process per engine in each invocation.
./scripts/run-battle-memory.sh
./scripts/run-battle-memory.sh
./scripts/run-battle-memory.sh
# Untimed validation of every reference and cache-enabled benchmark case.
cargo bench --locked --features ffi,match-cache --bench battle_bench -- --test