Chapter 13 · 2 July – 30 September
Roast, from 185 files to 1,424
Roast is Raku's official test suite, and the one definition of "correct" this project has ever used. On the first day Raku++ passed 185 of its files. Eleven weeks later v5.0.0 passed 1,423 of the 1,424 files Roast lists, and v5.1.0 passes all of them. The count did not rise steadily: it went down more than once, each time because the measurement had become more honest.
Two ways to count, and why the strict one leads
A file counts as passing only when every one of its tests passes. One stray failure in a 200-test file and the whole file is lost. The other figure is how many of all the tests the suite declares pass. The project reports both. The denominator for tests is "declared", not "ran": a file that dies before printing a result still contributes its planned count, recovered from the source, as failures.
That rule is why the early chart jumps. On 9 July the all-declared rate first read about 57%, against a denominator of 231,092. The next day the denominator was redefined to 187,714, and the same standing read about 73%. Part of that jump is a re-baseline, not progress. The dashboard marks it too.
When the number went down for the right reason
- Pair-form subtests ran for the first time.
subtest "name" => { … }is how most of Roast and most modules write subtests. It had never run its body: the block arrived as a Pair the builtin did not unpack, so every such subtest passed, empty. Fixing one line took the top line from 197,320 to 194,980, 2,340 fewer passes and 39 fewer passing files. v1.8.0's figures are marked "inflated" in the release notes ever since. - A file was given up on purpose.
integration/advent2012-day14.thad been passing on a guess that put 9 into a list of primes. - The harness ran what Roast lists. By default it now runs the files in Roast's own
spectest.data: 1,434 of them, not every.tin the checkout (1,464). Then Roast removed its eleven:P5files and added one, leaving 1,424. - A new headline. The lead figure became tests passing with Roast's own skip and todo markers left out, rather than counted as passes the way TAP counts them. Those markers cover 1,635 of the 220,055 tests (0.7%). The first reading under the new rule was 218,102 of 218,515 (99.81%).
The last stretch
Between v4.0.1 and v5.0.0 there were 249 commits, and most of them were this campaign. It ran in three shapes.
- Synopsis by synopsis. Bursts that each took one area to its ceiling: S02 literals and types, S03 operators (mostly list associativity), S13, S15, S19, S29, S32. Then the long tail, file by file, from the harness's
--failedlist. - Three tracks on the runtime model. Scalar containers: a variable bound to a second name shares one cell. A
gatherthat suspends instead of running ahead. One role per parameterization. Each track ran in its own worktree, in its own session. - Semantics sheets. Eighteen sheets of behaviour were extracted from Rakudo 2026.08, each rule with a probe and Rakudo's output. Eight of them are implemented.
One fix from that week shows the scale of what was left. S04-exceptions/catch.t runs a CATCH half a million times. Each iteration raised four C++ exceptions at about 20 µs each, so the file took 37 seconds and timed out. Once a die could reach its CATCH without a C++ throw, it took 0.4 seconds. Rakudo takes 2.5.
| reading | files | of |
|---|---|---|
| 21 Sep | 803 | 1,464 |
| 23 Sep | 818 | 1,464 |
| 25 Sep | 901 | 1,464 |
| 26 Sep | 1,189 | 1,434 |
| 27 Sep | 1,302 | 1,434 |
| 28 Sep | 1,415 | 1,434 |
| 29 Sep · v5.0.0 | 1,423 | 1,424 |
| 30 Sep · v5.1.0 | 1,424 | 1,424 |
The last file
v5.0.0's one failing test was test 3 of S16-io/eof.t, which reads .eof on a terminal. Roast marks it todo on macOS by release name ('Sonoma' | 'Sequoia' | 'Tahoe 26'), and the machine of record runs macOS 27. A day later Roast changed the marker to cover every macOS, and v5.1.0 passed all 1,424 files with no engine change for it.
The reference compiler, through the same harness
Rakudo 2026.08, run through the same harness on the same machine with Roast's own fudging applied, passes 1,414 of the same 1,424 files and 218,933 of 219,096 tests without skip and todo (99.93%). Six of the ten files it does not pass test regex features newer than that release. The comparison is not a race. It shows that the suite can be passed, and how the harness counts.
What it cost
The campaign was not free. v5.0.0 ran fib 25% slower than v4.0.1, and object construction 24% slower. The cost arrived in steps of 3 to 5 percent during 22–27 September. The final chapter is the work that won most of it back.