Chapter 2 · July – September

A regex engine of its own

Raku's grammars are its signature feature, and the regex engine under them was written from scratch: no library, no port. It was in the first commit. Over three months it learned to interpolate at match time, to rank alternatives the way the specification defines, and to parse a 2,653-line YAML file faster than the reference compiler.

62% → 85%of the regex synopsis's assertions, 2 to 10 July
52.7 s → 6.6 msloading the SPDX license list through JSON::Fast
0.46 sYAMLish on the course's table of contents; 0.86 s under Rakudo
The turnUntil August, an alternation | was ranked by a greedy probe that tried the alternatives. The specification asks for longest-token matching: rank each alternative by its declarative prefix, and never run user code while ranking. On 8 August a Thompson automaton over those prefixes replaced the probe as the default.

A matcher that interpolates

The generator behind the Raku course reads its table of contents through YAMLish, an indentation-sensitive YAML grammar that uses nearly every advanced regex feature at once. It sets lexical :my variables mid-match and uses them later as pattern atoms, evaluates code assertions against the live match, and takes quantifier bounds and subrule arguments at run time. Making it parse meant a matcher that interpolates at match time. By 1 August both XML and YAMLish passed their own zef test.

True longest-token matching

The probe could give answers Rakudo does not. "abcd" ~~ / [ ab {} cd ] | abc / matched abcd under the probe and abc under Rakudo, because the {} ends the declarative prefix. And my $n = 0; "x" ~~ / [ {$n++} y ] | x / left $n at 1: the probe had run user code just to rank.

A grammar as the benchmark

On 13 August the speed work took a real workload: a new, complete JSON parser written as a pure Raku grammar, profiled on real documents. Parsing, not the matcher, turned out to be under half the cost. About 44% went to building Match objects and running actions, and 38% to matching.

fileRaku++Rakudo
api.json323 ms520 ms
deep.json413 ms621 ms
strings.json214 ms470 ms
numbers.json187 ms285 ms

The same day found that one non-ASCII character made positional string operations quadratic. After the fix, JSON::Fast on api.json went from 52,820 ms to 335: 158 times faster.

JSON::Fast, native

JSON::Fast is the ecosystem's most depended-on module, and the one an interpreter serves worst: its parser walks text one grapheme at a time. Loading the 332 KB SPDX license list through it took 52.7 s. v3.0.0 ships an embedded version of the module: its own source, with only the parse machinery swapped for C++ at full fidelity (arbitrary-precision Int, exact Rat, surrogate pairs, strict escapes). Its own 14-file test suite is the gate, and passes. The SPDX parse takes 6.6 ms, against compiled Rakudo's 32.

YAMLish, in linear time

On v5.0.0 the course's 2,653-line table of contents did not finish parsing in 60 seconds. The fix-only v5.0.1 release, the same day, found two causes:

The parse now takes 0.46 s per process, against Rakudo's 0.86 s, and both produce the same data. Ten parses in one process take 4.51 s against 7.16 s.

A note on scale

At 18 MB of input, the pure-Raku grammar beats the hand-written JSON::Fast parser under Raku++ by 2× (16.2 s against 32.6 s). Under Rakudo the order is the other way round. Which code is fast depends on the engine running it.