Chapter 8 · 7 July – 28 September
Parallel by default
From the start, Raku++ ran start blocks on real OS threads, but only one thread at a time could interpret. A global interpreter lock kept the engine safe and kept a program on one core. On 9 August, v3.0.0 made parallel the default: the lock went, after a two-day campaign to make sure that a race in a program's data could no longer corrupt the interpreter.
race for ^400_000, once hyper for really fanned out- Plan A: safe points. Before this, a fire-and-forget compute
startheld the lock forever, so the program could never shut down. Workers became preemptible. The lock stayed. - The rules first. A written memory model, a stress suite of nine programs run in both modes with a ratchet (a new failure fails it, and so does a known failure that starts passing), and a ThreadSanitizer job in CI. The first run found the three gaps the plan had predicted:
atomicintlost updates, aChannelwith several producers hung, andSupplier.emitdropped emissions from other threads. - Registers per thread, caches decided once. Execution state moved into thread-local registers. The caches the interpreter writes into tree nodes became atomics that are decided once. ThreadSanitizer went from 35 reports to 6, then to none.
- Striped containers. Arrays and hashes take a lock stripe only while more than one worker is alive, so a single-threaded program pays nothing.
- Faster by default. The default flipped;
RAKUPP_GIL=1stayed for one release as the way back. On the Roast gate the parallel build timed out on 5 files, against 11 for the locked one.
A livelock in five minutes
S17-lowlevel/thread.t did not finish in five minutes. A sample of the stuck process showed cas running the user's block while holding its pool stripe. Three threads doing cas %seen{$_}, {.succ} each held one stripe and waited for another. The fix was a real compare-and-swap retry loop, and the file went from no output at all to 29 of 29.
What "N times faster" means
At v3.0.0 four workers ran 3.72 times as fast as one on the four performance cores. In September that was re-measured at 2.9×, and the parallel guide says why the ratio fell: the serial baseline had become about four times faster, and the threading had not. A single start block costs something (0.85× of plain code), and a contended atomic counter is slower in parallel than serial (0.79×). The guide gives the method for measuring your own case.
Later
- 27 Sep:
hyper forandrace forhad been parsed asdo forand run on the calling thread. They now fan out in batches of 64. The first version ran seven workers slower than one, because every worker bumped one shared reference count. Once fixed,race for ^400_000 { 1 }went from 145 ms to 44 ms (serial: 69 ms). This completed Roast's S07. - 28 Sep: a supply's queue had been written for the locked engine and had no lock of its own. One Roast test hung about once in 60 runs, and a looped reproduction lost 9 to 24 deliveries per run. It is locked now.