Chapter 13 · 2 July – 30 September

Roast, from 185 files to 1,424

Roast is Raku's official test suite, and the one definition of "correct" this project has ever used. On the first day Raku++ passed 185 of its files. Eleven weeks later v5.0.0 passed 1,423 of the 1,424 files Roast lists, and v5.1.0 passes all of them. The count did not rise steadily: it went down more than once, each time because the measurement had become more honest.

185Roast files passing completely on 2 July, the day of the first commit
90%of all declared tests, crossed by v1.0.0 on 22 July
803 → 1,423files, 21 to 29 September: the last stretch
1,424 / 1,424files in v5.1.0; 218,420 of 218,420 tests with skip and todo left out
Roast, first commit to now. Blue is the share of files in which every test passes; green is the share of all declared tests that pass. Hover for each release.
The turnFor eleven weeks the file count rose by about one file a day, against a moving denominator. In the last nine days it went from 803 to 1,423. That was not a new engine. The work was organised differently: by synopsis, then by runtime model, then file by file from the harness's list of failures.

Two ways to count, and why the strict one leads

A file counts as passing only when every one of its tests passes. One stray failure in a 200-test file and the whole file is lost. The other figure is how many of all the tests the suite declares pass. The project reports both. The denominator for tests is "declared", not "ran": a file that dies before printing a result still contributes its planned count, recovered from the source, as failures.

That rule is why the early chart jumps. On 9 July the all-declared rate first read about 57%, against a denominator of 231,092. The next day the denominator was redefined to 187,714, and the same standing read about 73%. Part of that jump is a re-baseline, not progress. The dashboard marks it too.

When the number went down for the right reason

The last stretch

Between v4.0.1 and v5.0.0 there were 249 commits, and most of them were this campaign. It ran in three shapes.

One fix from that week shows the scale of what was left. S04-exceptions/catch.t runs a CATCH half a million times. Each iteration raised four C++ exceptions at about 20 µs each, so the file took 37 seconds and timed out. Once a die could reach its CATCH without a C++ throw, it took 0.4 seconds. Rakudo takes 2.5.

readingfilesof
21 Sep8031,464
23 Sep8181,464
25 Sep9011,464
26 Sep1,1891,434
27 Sep1,3021,434
28 Sep1,4151,434
29 Sep · v5.0.01,4231,424
30 Sep · v5.1.01,4241,424

The last file

v5.0.0's one failing test was test 3 of S16-io/eof.t, which reads .eof on a terminal. Roast marks it todo on macOS by release name ('Sonoma' | 'Sequoia' | 'Tahoe 26'), and the machine of record runs macOS 27. A day later Roast changed the marker to cover every macOS, and v5.1.0 passed all 1,424 files with no engine change for it.

The reference compiler, through the same harness

Rakudo 2026.08, run through the same harness on the same machine with Roast's own fudging applied, passes 1,414 of the same 1,424 files and 218,933 of 219,096 tests without skip and todo (99.93%). Six of the ten files it does not pass test regex features newer than that release. The comparison is not a race. It shows that the suite can be passed, and how the harness counts.

What it cost

The campaign was not free. v5.0.0 ran fib 25% slower than v4.0.1, and object construction 24% slower. The cost arrived in steps of 3 to 5 percent during 22–27 September. The final chapter is the work that won most of it back.