Chapter 9 · 9 August – 29 September
Honest instruments
Every number on these pages comes from a tool: the Roast harness, the module battery, the benchmark runner, the documentation sweep. In late August the tools themselves were put under review, and what that found changed how every figure since has been taken.
7release gates in v3.21.0, all green over a hash function returning wrong answers
8 / 8planted defects caught once every gate had to prove it could fail
25more defects found by reviewing the instruments in v3.23.0
The turnBefore v3.22.0 a green gate was taken on trust. From v3.22.0 on, a gate counts only after it has been shown to go red on a defect planted for it. The gate is always the list of files that pass, never the count.
- The procedure, run. v3.0.0 had been tagged without running the written release procedure. Running it produced lower figures, and they were published: 197,080 tests where v3.0.0 had claimed 197,191, and 12–16 timeouts a run where the notes had claimed 5. It also found that three of the five platform binaries came from the wrong commit.
- A figure corrected in the open. The release-day tally double-counted an interleaved log and published 97 of 105 for the module battery. The measured figure was 48 of 59, two below the release before. The notes now say so.
- All seven gates passed, and the release shipped a wrong answer.
Digest's RIPEMD returned incorrect hashes for every input. Its constant table was built with 5 entries instead of 80. One gate of seven saw it, and the first diagnosis blamed.flat. Checked against Rakudo, that was wrong: the fault was in binding. - The instruments, fixed and then proved. Six of the seven gates had a defect of their own. Each was fixed, and then each was given a planted fault to catch.
- The re-baseline. Reviewing the tools before re-measuring through them found twenty-five more defects, three of which would have corrupted that release's own numbers. It also found that "8 of 8" had been a correct count but an incorrect claim: the conformance gate had no red path at all.
- The battery had been comparing Raku++ with itself. The battery runs each distribution's suite under both engines. It started its reference engine by the name
raku, and on the machine of record that name had meant Raku++ since 12 September. Against Rakudo, the release candidate had broken eleven distributions. Ten were fixed before the tag.
What the instruments got wrong
- The harness corrupted its own status lines. A child process's stderr (14,026 lines of it) was spliced through four
[PASS]lines, so four files a run lost their path and fell out of the list. The list is now written as data, not scraped from a log. - A fork storm that passed. A compiled binary under test re-ran itself: 1,253 processes at load average 450, and the gate reported PASSED.
- A comparison against itself, in miniature. The battery compared a run against the file it had just written, so it reported "no regression" while
Digestwent from 4 of 4 to 3 of 4. - A dashboard that published 638 while the tag said 643. Bare table cells were read by a script that could not tell which number was which. A figure checker now reads the release notes the way a person would.
How a figure is taken now
- The list, not the count. The gate is the sorted list of files that pass completely, and it must have no removals. The count is the headline; the list is the check.
- Say what is being measured. Three different
rakuppbinaries once answered on the same PATH. Every run now records which binary and which Rakudo it used. - Run alone. Timing-sensitive files time out on a loaded machine. Gates run with nothing else beside them, and a release quotes the profile that repeats across runs, not the best one. In v3.7.0, five passes each dropped a different file to a timeout; the union of their lists was the figure, and the diff against the release before was empty.
- Interleave, then compare. A performance A/B runs the old and new binaries in turn (A, B, A, B…) on the same machine, so both see the same conditions.