Skip to content

PR 173: Full Oniguruma Comparison

This report evaluates general regex tasks and scanner workloads for PR #173, the opt-in match-cache reconstruction. The initial reference run covers 72 timings across 29 paired Ferroni/C scenarios, with the regex crate in 14 compatible cases. A follow-up adds seven everyday regex tasks measured with all three engines. Separate runs measure the opt-in cache on three documents and compare process memory. Syntax highlighting is one integration workload; it does not stand in for the general engine.

The general-regex follow-up and initial reference run both compile the new failed-state match cache out. Its benefit on pathological inputs and review requirements are documented in the match-cache evaluation. The separate cache-on/off document table measures the cost of opting in on ordinary documents. It does not attribute existing Ferroni advantages to the new cache.

What the comparison covers

The primary question is how Ferroni performs against the C engine whose behavior it ports: Oniguruma. The main tables compare those two engines, including patterns that use Oniguruma features unavailable in regex.

Capability exercised hereFerroniOnigurumaRust regex
Literals, classes, quantifiers, alternationSupportedSupportedSupported
Capture groups and the tested Unicode patternsSupportedSupportedSupported
Lookahead and lookbehindSupported and measuredSupported and measuredNot supported
BackreferencesSupported and measuredSupported and measuredNot supported
RegSet/scanner workload in this harnessMeasuredMeasuredNo equivalent path measured

This is a scope map for the benchmarks, not a complete compatibility matrix. regex deliberately excludes lookarounds and backreferences. The seven everyday tasks below all use shared syntax; they do not exercise those differences. A result on that subset cannot establish interchangeability for an existing Oniguruma pattern set.

The shared-syntax appendix retains every regex measurement as context for applications whose patterns fit that subset. Unsupported syntax and an unmeasured API path are separate limitations; neither is a timing result. There is no winner count across engines with different feature coverage.

Everyday regex tasks: Ferroni vs Oniguruma

This follow-up focuses on the general engine: three validation batches and four complete text-processing tasks, with no TextMate scanner. Ferroni has the lower point estimate than C in all seven cases, by 1.08–1.98x. The scanner speedups are not a useful estimate for ordinary regex work.

Task (whole batch/text)FerroniOnigurumaC / Ferroni
Email shape, 64 inputs6.722 us12.843 us1.91x
UUID shape, 64 inputs11.079 us15.773 us1.42x
Number syntax, 64 inputs8.034 us14.401 us1.79x
Access log, 64 records / 6 fields32.410 us34.866 us1.08x
Extract 64 URLs9.871 us19.546 us1.98x
Extract 384 Unicode words63.527 us81.591 us1.28x
Redact 128 email addresses78.465 us87.898 us1.12x

Times cover the whole batch or text, not one match. Bold marks the lowest point estimate between Ferroni and C, not statistical significance. The validators repeat eight labeled inputs eight times: 32 accepted and 32 rejected, including early rejections and failures near the end. They measure boolean results without capture output. Email and UUID are simplified shape checks, not complete RFC or UUID-version validation; the number case checks signed decimal/exponent syntax without numeric conversion.

The text tasks use deterministic synthetic inputs from general_regex.rs:

  • Access log: 5,964 bytes, 64 matching records, six captured fields each, plus 16 nonmatching messages.
  • URL extraction: 5,196 bytes and 64 HTTP/HTTPS links amid plain text.
  • Unicode words: 3,648 bytes and 384 words using Latin, Greek, Cyrillic, Arabic, Han characters, and a decomposed accent. The pattern is \p{L}[\p{L}\p{M}]*.
  • Email redaction: 5,218 bytes, 128 addresses, and similar-looking nonmatches. Every match becomes [email] through the same output builder for all engines.

Each text operation collects every match and capture into byte-offset vectors. Capture scratch and result allocation are inside timing; the region or capture-location storage is reused between searches within one operation. The redaction row includes the common output builder and is not a comparison of native replacement APIs. Compilation, input generation, and correctness checks are outside timing. Patterns are compiled once and reused. These inputs contain no empty matches; empty-match iteration policy is outside this scope.

Before timing, all three engines agree on each labeled validation result and every extracted match/capture boundary. The preflight also checks expected match counts, concrete log-field values, and the exact redacted output. The benchmark smoke run covers all 96 cases with ffi,match-cache enabled.

Access-log capture and URL extraction were repeated separately with the same configuration. Ferroni remained faster than C on these two inputs. The appendix retains the third engine's repeat timings as shared-syntax context.

Repeat (whole text)FerroniOnigurumaC / Ferroni
Access log31.591 us33.240 us1.05x
URL extraction10.080 us19.530 us1.94x

The main table retains the original run. These are a few representative tasks, not an application-traffic distribution or a cross-platform engine ranking. Backreferences, lookarounds, single-pattern compilation, and long no-match searches remain in the initial reference tables below; syntax highlighting remains a separate integration workload.

The follow-up was measured on 2026-09-26 at harness commit 03a6ba69c5be294e77ca20ce9f0a08db4973f0a4, with the same machine, compilers, dependencies, input pins, and 30-sample / 500 ms warmup / 4 s measurement settings listed below. Engine sources are unchanged from 9f18790. It uses --features ffi, so the opt-in match cache is compiled out. Each case runs Ferroni, then C, then regex; the repeats retain that fixed order. This is a separate run from the original scanner/feature suite.

The follow-up raw artifact contains all 21 initial measurements and six repeat measurements, 95% confidence intervals, individual samples, exact patterns, workload sizes, and source hashes.

cargo bench --locked --features ffi --bench battle_bench -- \
  general_regex --noplot --save-baseline pr173-general-regex
# Render before repeats overwrite Criterion's `new` directories.
python3 scripts/gen_battle_tables.py --general-only

cargo bench --locked --features ffi --bench battle_bench -- \
  access_log_captures --noplot --save-baseline pr173-general-log-repeat
cargo bench --locked --features ffi --bench battle_bench -- \
  url_extraction --noplot --save-baseline pr173-general-url-repeat

Initial reference-run measurement context

FieldValue
Date2026-09-26
Ferroni/harness commit719a51f9e0caba622e90470f493f5f5b89848d64
Engine source commit9f1879007d735e791fc33c69133920ff22fc1640 (unchanged by the harness update)
HostMac13,2, Apple M1 Ultra, 64 GiB RAM
OS / targetmacOS 27.0 (26A428), aarch64-apple-darwin
Rustrustc 1.96.0 (ac68faa20), LLVM 22.1.2
C compilerApple clang 21.0.0 (clang-2100.3.34.2)
ProfilesOptimized Rust bench profile, thin LTO; C Oniguruma at -O3
DependenciesCriterion 0.8.2; regex 1.13.1; checked-in Cargo.lock
Criterion30 samples, 500 ms warmup, 4 s measurement per case
Onigurumaf95747b462de672b6f8dbdeb478245ddf061ca53
Full grammarsTypeScript: 279 patterns; CSS: 117; Rust: 81

Criterion extends slow cases when four seconds would not collect 30 samples. Tables report the slope estimate when available, otherwise the mean. The raw artifact contains all 83 timing measurements (72 reference, two CSS repeat, nine cache comparison), 95% confidence intervals, individual sample durations/iteration counts, all three memory runs, source hashes, and the input pins from battle_inputs.toml.

The full run, repeat, cache comparison, and memory processes ran sequentially. This is a developer-workstation measurement with a fixed engine order, not a cross-platform or randomized campaign. Small percentage differences require repetition before becoming release criteria.

Equivalent work and correctness checks

The harness previously requested capture output from C in timed text and single-pattern searches while Ferroni received no output region. This run requests no capture output from either engine. Separate, untimed assertions compare exact match positions and raw capture bounds. Scanner timings include captures for both engines; their preflight compares pattern indices and every capture boundary along the complete fixture trace. These checks run in optimized builds too, including the cache-enabled lane.

The scanner fixtures are ASCII, so their byte and UTF-16 offsets coincide. This validates these benchmark inputs; it does not replace the broader Unicode, differential, and fuzz evidence in the cache evaluation. regex is syntax-limited context, and its find API returns both start and end positions.

The older published benchmark tables remain a dated snapshot on a different machine and toolchain, with the previous harness. Do not calculate a version-to-version regression from the two pages.

Initial reference suite

In the initial run, default Ferroni has a lower point estimate than C Oniguruma in all 23 search/scanner scenarios and in three of six compilation scenarios. Whole-document tokenization is approximately 3.0x (TypeScript), 31.7x (CSS), and 10.3x (Rust) faster in the full run. CSS has a noisy initial sample; the targeted repeat and both sets of raw samples are retained below. These ratios describe this suite, not a universal speedup.

Oniguruma still compiles the simple literal, lookbehind, and full Rust grammar faster. Shared-syntax timings for regex are retained in the appendix rather than treated as a ranking of engines with equivalent capabilities.

Bold marks the lower Ferroni/C point estimate, not statistical significance. Factors are C time divided by Ferroni time; values above one favor Ferroni.

Text search and log scanning

ScenarioFerroniOnigurumaC / Ferroni
Literal in 50 KB70.279 ns162.110 ns2.3x
No match, 50 KB1.554 us9.434 us6.1x
No match, 10 KB361.350 ns1.881 us5.2x
Field extract, 50 KB94.270 ns172.352 ns1.83x
Timestamp, 50 KB95.618 ns192.295 ns2.0x
RegSet multi-pattern (5)100.562 ns399.525 ns4.0x

Pattern matching

CategoryFerroniOnigurumaC / Ferroni
Literal exact102.612 ns157.929 ns1.54x
Quantifier greedy62.308 ns200.063 ns3.2x
Lookaround combined87.678 ns248.714 ns2.8x
Unicode \p{Greek}+103.399 ns213.873 ns2.1x
Backref (\w+) \184.129 ns174.816 ns2.1x
Case-insensitive phrase104.720 ns185.957 ns1.78x
Alternation, 2 branches66.211 ns160.676 ns2.4x
Alternation, 10 branches55.980 ns242.930 ns4.3x
Named capture date221.172 ns260.521 ns1.18x

Compilation

PatternFerroniOnigurumaC / Ferroni
Literal507.053 ns493.223 ns0.97x
Named capture4.113 us6.337 us1.54x
Lookbehind1.161 us611.973 ns0.53x

Scanner with full Shiki TextMate grammars

ScenarioFerroniOnigurumaFactor
TypeScript (279 patterns)
Compile11.088 ms17.454 ms1.57x
First match, short line85.084 ns25.569 us300.5x
Tokenize full line2.707 us204.592 us75.6x
CSS (117 patterns)
Compile14.870 ms18.541 ms1.25x
Tokenize (multi-line)221.197 us14.712 ms66.5x
Rust (81 patterns)
Compile319.448 us195.092 us0.61x
First match172.934 ns5.505 us31.8x
Tokenize full line4.930 us80.209 us16.3x

Scanner on whole documents, line by line

DocumentFerroniOnigurumaFactor
TypeScript (279 patterns), 28 lines1.035 ms3.129 ms3.0x
CSS (117 patterns), 19 lines100.494 us3.190 ms31.7x
Rust (81 patterns), 31 lines105.659 us1.086 ms10.3x

The scanner-highlighting rows reuse the same OnigString. Ferroni's existing per-string memo avoids repeated fallback work, while the C scanner wrapper consults its per-pattern cache only at 1,000 bytes or more. All those fixtures are shorter. The large warm-path ratios therefore describe these two wrapper policies as well as the engines.

The whole-document rows tokenize each fixture line by line, preserving trailing newlines and changing string identity per line. Compilation is excluded, and the per-line strings are prepared before timing. These are scanner workloads over full grammar pattern sets, not end-to-end Shiki/TextMate rendering with grammar-state transitions.

The initial CSS Ferroni estimate was 100.494 us, with a 95% confidence interval of 90.240–115.971 us. A targeted repeat with the same settings measured 89.059 us for Ferroni and 3.126 ms for C (35.1x), with a much narrower Ferroni interval of 88.933–89.158 us. The main table deliberately retains the complete run instead of substituting the faster repeat. No single aggregate combines cold, warm, search, and compilation ratios.

Opt-in cache on ordinary documents

This separate ffi,match-cache build measures runtime cache off, runtime cache on (default 16 MiB budget and adaptive activation), and C in the same executable. The inputs, trace validation, and timing settings match the document rows above. Opt-in cost is (cache on / cache off - 1) within this feature-enabled build; do not calculate that cost across different executables.

DocumentCache offCache onOnigurumaOpt-in cost
TypeScript1.126 ms1.141 ms3.162 ms+1.38%
CSS99.183 us102.197 us3.288 ms+3.04%
Rust120.742 us128.018 us1.113 ms+6.03%

The cache protects eligible expensive failures; ordinary grammar documents still pay bookkeeping costs. This run does not justify default enablement. The feature-disabled release comparison and pathological cases remain in the separate evaluation.

Process memory

The memory harness compiles the same full TypeScript grammar and scans a deterministic 1,204,992-byte source containing 16,002 lines. It requests the first match from each line, not complete tokenization. Both engines reported 16,002 scanned and matched lines in each of three runs. Each engine runs in its own process; Ferroni uses the default build without match-cache.

EnginePhaseMedian peak RSSRange (three runs)
FerroniCompile15,106,048 B15,089,664–15,122,432 B
FerroniCompile + line scan15,204,352 B15,187,968–15,220,736 B
OnigurumaCompile15,728,640 B15,482,880–15,990,784 B
OnigurumaCompile + line scan15,876,096 B15,679,488–16,187,392 B

These are process high-water RSS values from getrusage, including the driver, input, grammar storage, allocator, and compiled engine. The scan value includes the preceding compile phase. They are not engine heap sizes, incremental scan allocations, or measurements of the opt-in cache's 16 MiB budget. See the memory methodology for the harness boundaries.

Shared-syntax context only: regex

These measurements answer a narrower question: how quickly each engine runs the particular patterns supported by all three. They do not compare the full Oniguruma feature set or establish whether regex can replace Ferroni. regex excludes lookarounds and backreferences as part of its design. Its timings remain useful when an application's patterns fit that subset. Capture groups and Unicode are supported; those features alone do not require an Oniguruma-compatible engine.

The original measurements are retained below without winner counts or an overall ranking. Only measured shared-syntax cases appear here. Lookarounds, backreferences, and the unmeasured RegSet/scanner path are excluded explicitly, not assigned a slow or zero result. Times and measurement conditions are unchanged from their corresponding main tables.

Shared syntax: everyday regex tasks

Task (whole batch/text)FerroniOnigurumaregex (shared syntax only)
Email shape, 64 inputs6.722 us12.843 us1.522 us
UUID shape, 64 inputs11.079 us15.773 us1.769 us
Number syntax, 64 inputs8.034 us14.401 us1.033 us
Access log, 64 records / 6 fields32.410 us34.866 us35.726 us
Extract 64 URLs9.871 us19.546 us13.364 us
Extract 384 Unicode words63.527 us81.591 us32.798 us
Redact 128 email addresses78.465 us87.898 us13.915 us

Shared syntax: everyday task repeats

Repeat (whole text)FerroniOnigurumaregex (shared syntax only)
Access log31.591 us33.240 us35.684 us
URL extraction10.080 us19.530 us13.377 us

Shared syntax: text search and log scanning

ScenarioFerroniOnigurumaregex (shared syntax only)
Literal in 50 KB70.279 ns162.110 ns9.757 ns
No match, 50 KB1.554 us9.434 us1.459 us
No match, 10 KB361.350 ns1.881 us296.226 ns
Field extract, 50 KB94.270 ns172.352 ns54.945 ns
Timestamp, 50 KB95.618 ns192.295 ns63.992 ns

Shared syntax: pattern matching

CategoryFerroniOnigurumaregex (shared syntax only)
Literal exact102.612 ns157.929 ns11.091 ns
Quantifier greedy62.308 ns200.063 ns64.364 ns
Unicode \p{Greek}+103.399 ns213.873 ns58.792 ns
Case-insensitive phrase104.720 ns185.957 ns60.377 ns
Alternation, 2 branches66.211 ns160.676 ns45.972 ns
Alternation, 10 branches55.980 ns242.930 ns20.073 ns
Named capture date221.172 ns260.521 ns44.633 ns

Shared syntax: compilation

PatternFerroniOnigurumaregex (shared syntax only)
Literal507.053 ns493.223 ns2.896 us
Named capture4.113 us6.337 us208.787 us

Reproducing this snapshot

To reproduce the original 72-measurement suite exactly, use harness commit 719a51f. The current full suite also includes general_regex; the follow-up commands above isolate that addition. Run each timing command to completion before starting the next one. Group configuration in battle_bench.rs fixes the Criterion settings listed above.

./scripts/prepare-oniguruma-sources.sh
cargo bench --locked --features ffi --bench battle_bench -- \
  --noplot --save-baseline pr173-reference
# Render before the filtered runs overwrite Criterion's `new` directories.
python3 scripts/gen_battle_tables.py

cargo bench --locked --features ffi --bench battle_bench -- \
  scanner_documents/css_117 --noplot --save-baseline pr173-css-repeat
cargo bench --locked --features ffi,match-cache --bench battle_bench -- \
  scanner_documents --noplot --save-baseline pr173-cache-documents

# Three sequential invocations, one process per engine in each invocation.
./scripts/run-battle-memory.sh
./scripts/run-battle-memory.sh
./scripts/run-battle-memory.sh

# Untimed validation of every reference and cache-enabled benchmark case.
cargo bench --locked --features ffi,match-cache --bench battle_bench -- --test