Chapter 10 · 9 August – 30 September

Smaller values, real frames

In August the question changed. Instead of "what is slow in this program", it became "what does every operation pay". The answer came from reading how Perl 5 does it. In two days a value went from 344 bytes to 128, variables stopped being looked up by name, and storing a result stopped costing more than computing it.

344 → 128bytes per value, 21–22 August
−34%loopsum, when variables became slots in a frame
28 vs 108 nsthe cost of an add, against the cost of storing its result
4.60 → 0.03 sfor ^200000 { next if $_ % 2 }
loopsum, every release, one machine. Blue is Raku++ interpreting. The dashed line is the Rakudo that was current on each release's day.
The turnA remark that natively compiled Raku++ was "only twice slower than Perl 5" became a benchmark kernel, and then a reading of five Perl 5 source files. The notes list eight techniques. A Perl scalar is 24 bytes, and ours was 344. Perl resolves each lexical to an array offset at compile time. Every value-producing op writes into a pre-assigned target slot. Most of what followed was those items, one plan each, each with a stated way to be proved wrong.

The value

The target had been 64 bytes. A September study showed why the work stopped at 128. The last 24 bytes are all string. Moving the string out of the value made short strings slower (0.60×). Keeping it inline would change the type at about 1,549 sites. Reaching 16 bytes would mean values stop being reference-counted, which is a rewrite of assignment and binding, not a batch. A 56-byte design is priced and waiting.

Hashes like Perl's

On 21 August std::map gave way to an open-addressed table that stores each key's hash, as Perl's does: the hash kernel −30%, the hash-fill kernel −19% interpreted and −24% compiled. On that kernel the compiled binary passed Perl: 78 ms against 82.

Pads

Until 22 August every variable read hashed the variable's name into the current scope's map and walked up the parent scopes on a miss. Perl resolves a lexical to an array offset at compile time, and reading it is one indexed load. Raku++ now does the same.

The first A/B was flat on the assignment kernel, which was exactly what the plan said would prove it wrong. A sample showed the cost had moved to creating a scope on every iteration, so flat loop bodies now keep one. After that: assignment −25%, loopsum −34%, hash −24%, strcat −24%. The same batch fixed a recursion bug in is rw that the Forth showcase had found. Deriving the frame instead of tracking it in a register fixed some 109 re-pointing sites at once.

The cost of a result

The plan priced one loop iteration: a floor of 65 ns, reading a variable 37, the add itself 28, and storing its result 108. The arithmetic cost 28 ns, and storing its result cost 108. The first slice of Perl's target-slot idea, a lane for simple assignment, took the assignment kernel down 40%. Later slices took sub calls and string passing down 19.5% each.

Levers, one plan each

A correction worth keeping

On 22 August the benchmarks page said the interpreter had overtaken Rakudo on fib. The Rakudo in that comparison had been running under Rosetta 2, emulated. Re-measured natively, fib was still behind (369.8 ms against 311.4 at v3.7.0). The one-machine re-measurement of every release, two chapters on, exists because of that mistake.