Chapter 2 · July – September
A regex engine of its own
Raku's grammars are its signature feature, and the regex engine under them was written from scratch: no library, no port. It was in the first commit. Over three months it learned to interpolate at match time, to rank alternatives the way the specification defines, and to parse a 2,653-line YAML file faster than the reference compiler.
| was ranked by a greedy probe that tried the alternatives. The specification asks for longest-token matching: rank each alternative by its declarative prefix, and never run user code while ranking. On 8 August a Thompson automaton over those prefixes replaced the probe as the default.A matcher that interpolates
The generator behind the Raku course reads its table of contents through YAMLish, an indentation-sensitive YAML grammar that uses nearly every advanced regex feature at once. It sets lexical :my variables mid-match and uses them later as pattern atoms, evaluates code assertions against the live match, and takes quantifier bounds and subrule arguments at run time. Making it parse meant a matcher that interpolates at match time. By 1 August both XML and YAMLish passed their own zef test.
True longest-token matching
The probe could give answers Rakudo does not. "abcd" ~~ / [ ab {} cd ] | abc / matched abcd under the probe and abc under Rakudo, because the {} ends the declarative prefix. And my $n = 0; "x" ~~ / [ {$n++} y ] | x / left $n at 1: the probe had run user code just to rank.
- The plan, one of the three pillars of v3.0.0. Baseline: a 60-key JSON grammar parsed 20 times took 31 ms here and 147 ms under Rakudo.
- A Thompson NFA over each alternative's declarative prefix (
src/LtmNfa.cpp). Behind a flag, it beat the probe on the full Roast run with a strict subset of its failures. The same night it became the default:protoregex.t30 of 30, the grammar benchmark unchanged at 28 ms.
A grammar as the benchmark
On 13 August the speed work took a real workload: a new, complete JSON parser written as a pure Raku grammar, profiled on real documents. Parsing, not the matcher, turned out to be under half the cost. About 44% went to building Match objects and running actions, and 38% to matching.
| file | Raku++ | Rakudo |
|---|---|---|
| api.json | 323 ms | 520 ms |
| deep.json | 413 ms | 621 ms |
| strings.json | 214 ms | 470 ms |
| numbers.json | 187 ms | 285 ms |
The same day found that one non-ASCII character made positional string operations quadratic. After the fix, JSON::Fast on api.json went from 52,820 ms to 335: 158 times faster.
JSON::Fast, native
JSON::Fast is the ecosystem's most depended-on module, and the one an interpreter serves worst: its parser walks text one grapheme at a time. Loading the 332 KB SPDX license list through it took 52.7 s. v3.0.0 ships an embedded version of the module: its own source, with only the parse machinery swapped for C++ at full fidelity (arbitrary-precision Int, exact Rat, surrogate pairs, strict escapes). Its own 14-file test suite is the gate, and passes. The SPDX parse takes 6.6 ms, against compiled Rakudo's 32.
YAMLish, in linear time
On v5.0.0 the course's 2,653-line table of contents did not finish parsing in 60 seconds. The fix-only v5.0.1 release, the same day, found two causes:
- An alias rebuilt its subtree once per key.
<alias=rule>built its Match once for each name, so nested aliases rebuilt every subtree 2depth times. A list four levels deep took 70 s. - A one-character lookbehind scanned from the start of the input.
m:g/<!after <.alpha>> a/over 40,000 words took 70 s; it takes 0.02 s now.
The parse now takes 0.46 s per process, against Rakudo's 0.86 s, and both produce the same data. Ten parses in one process take 4.51 s against 7.16 s.
A note on scale
At 18 MB of input, the pure-Raku grammar beats the hand-written JSON::Fast parser under Raku++ by 2× (16.2 s against 32.6 s). Under Rakudo the order is the other way round. Which code is fast depends on the engine running it.