Chapter 8 · 7 July – 28 September

Parallel by default

From the start, Raku++ ran start blocks on real OS threads, but only one thread at a time could interpret. A global interpreter lock kept the engine safe and kept a program on one core. On 9 August, v3.0.0 made parallel the default: the lock went, after a two-day campaign to make sure that a race in a program's data could no longer corrupt the interpreter.

35 → 0ThreadSanitizer reports, in two batches
26 / 26stress programs clean, in both modes
2.9×four workers against one, contention-free, re-measured in September
145 → 44 msrace for ^400_000, once hyper for really fanned out
The turnThe design record weighed three options: lock everything finely, impose an ownership discipline on the program, or run isolated interpreters. It chose to harden the runtime rather than every user structure. Unsynchronised sharing is the program's bug, but a race in the program's data must never corrupt the interpreter.

A livelock in five minutes

S17-lowlevel/thread.t did not finish in five minutes. A sample of the stuck process showed cas running the user's block while holding its pool stripe. Three threads doing cas %seen{$_}, {.succ} each held one stripe and waited for another. The fix was a real compare-and-swap retry loop, and the file went from no output at all to 29 of 29.

What "N times faster" means

At v3.0.0 four workers ran 3.72 times as fast as one on the four performance cores. In September that was re-measured at 2.9×, and the parallel guide says why the ratio fell: the serial baseline had become about four times faster, and the threading had not. A single start block costs something (0.85× of plain code), and a contended atomic counter is slower in parallel than serial (0.79×). The guide gives the method for measuring your own case.

Later