Chapter 7 · 11 July – 1 August
Slow programs, and what they pointed at
For its first month, Raku++ got faster because some small program felt slow. The example programs, Mandelbrot, N-Queens, Life and Brainfuck, were timed against Rakudo, and each slow one pointed at a single engine fault. In late July the work turned to profiles, and a benchmark gate began to decide whether a release could ship.
next stopped being a C++ exceptionfib in v1.5.1, from a profile-ranked campaign$a OP $b in a loop, from node specialization (v1.7.0)What the slow programs pointed at
- Control flow was exceptions.
return,next,lastandredowere C++ exceptions. N-Queens at N=10 threw about half a million times and spent some 90% of its profile unwinding. They became a thread-local flag with a per-activation frame counter; only a jump that crosses a closure or a builtin's callback still throws. N=8 went from 0.69 s to 0.09 s, N=10 from 16.0 s to 1.81 s, and Life from 2.18 s to 1.31 s. - Rats re-reduced. The Mandelbrot demo went from 6.83 s to 0.32 s. The numbers chapter tells that one.
- Plain variables skipped the special cases, and loops reused one scope per iteration: loopsum −24%, fib −22%, hash −22%.
- An array read copied the value three times.
@a[$i]copied a value of about 300 bytes on every read. The copy constructor was 98% of Brainfuck's profile. A pointer fast path took it from 0.29 s to 0.01 s. - A function object per iteration. A
std::functionallocated on every loop iteration: loopsum 227 → 190 ms. The benchmark gate gained its loop and hash kernels.
Decided once
v1.5.0 (29 July) found fib walking its own body about 2.7 million times, through an allocating callback, to decide whether any declaration needed hoisting. The answer never changes, so it is now decided once per block and stored on the tree node. fib went from 911 to 831 ms. That became a pattern: a question whose answer depends only on the program is answered once and remembered on the node. When the engine went parallel, those answers became atomics that are decided once (the parallel chapter).
A campaign ranked by the profile
v1.5.1 changed no behaviour at all; Roast was identical byte for byte. Five candidates were ranked from a profile. Three landed, one was measured and dropped, and one was measured and never attempted.
- A call's argument vector is moved, not copied: −9% on call-heavy code, from four lines. It was found because the call path accounted for 514 of about 580 allocation samples.
- The invocant is passed by reference, not copied to serve four arms out of 352: −3.4%.
- Rarely used scope containers are created on demand: a scope went from 256 to 72 bytes, −2.4%.
fib: 903.3 → 744.1 ms. Slot-indexed variables were billed as the biggest architectural win of all, but measured a ceiling of about 4% and were set aside. Three weeks later, with the rest of the frame work done, they took loopsum down 34% (the next speed chapter).
Shapes
v1.7.0 (1 August) taught the two hottest evaluators to recognise four syntactic shapes: $var OP literal, literal OP $var, $var OP $var, and @arr[$var] or @arr[literal]. The verdict is recorded on the node, and the path skips what the general case must do. $a OP $b −18.3%, $a OP 1 −17.7%, fib −11.7%, @a[$i] −8.7%.
Measured and not done
- Constant folding found 37 foldable sites in 51,353 nodes across 48 real programs.
- A hash map for method dispatch: 56% of the arms dispatch on the invocant's type, and one lookup costs about as much as 19 of the name compares it would replace.
- A register IR instead of the tree (8 August): dispatch is under 1% of a node visit. Re-measured in September, dispatch cost 0.32 ns against a node visit of 46 to 85 ns.
- The gate cried wolf once. fib read +27% because background processes had saturated the machine. Since then the old and new binaries are interleaved.