Chapter 10 · 9 August – 30 September
Smaller values, real frames
In August the question changed. Instead of "what is slow in this program", it became "what does every operation pay". The answer came from reading how Perl 5 does it. In two days a value went from 344 bytes to 128, variables stopped being looked up by name, and storing a result stopped costing more than computing it.
for ^200000 { next if $_ % 2 }The value
- 376 bytes: five strings and eleven shared pointers in every value.
- 344: the type and container tags interned.
- 200: a census found that 25.6 million of about 30 million destroyed values had none of the eleven pointers set. The rarely used fields moved to a cold block allocated on demand.
- 128: five payload pointers became one tagged slot.
The target had been 64 bytes. A September study showed why the work stopped at 128. The last 24 bytes are all string. Moving the string out of the value made short strings slower (0.60×). Keeping it inline would change the type at about 1,549 sites. Reaching 16 bytes would mean values stop being reference-counted, which is a rewrite of assignment and binding, not a batch. A 56-byte design is priced and waiting.
Hashes like Perl's
On 21 August std::map gave way to an open-addressed table that stores each key's hash, as Perl's does: the hash kernel −30%, the hash-fill kernel −19% interpreted and −24% compiled. On that kernel the compiled binary passed Perl: 78 ms against 82.
Pads
Until 22 August every variable read hashed the variable's name into the current scope's map and walked up the parent scopes on a miss. Perl resolves a lexical to an array offset at compile time, and reading it is one indexed load. Raku++ now does the same.
The first A/B was flat on the assignment kernel, which was exactly what the plan said would prove it wrong. A sample showed the cost had moved to creating a scope on every iteration, so flat loop bodies now keep one. After that: assignment −25%, loopsum −34%, hash −24%, strcat −24%. The same batch fixed a recursion bug in is rw that the Forth showcase had found. Deriving the frame instead of tracking it in a register fixed some 109 re-pointing sites at once.
The cost of a result
The plan priced one loop iteration: a floor of 65 ns, reading a variable 37, the add itself 28, and storing its result 108. The arithmetic cost 28 ns, and storing its result cost 108. The first slice of Perl's target-slot idea, a lane for simple assignment, took the assignment kernel down 40%. Later slices took sub calls and string passing down 19.5% each.
Levers, one plan each
- Methods stop allocating. An issue timed a graph library at 4× Rakudo. A user method call was 5.8× Rakudo's cost and a multi method 10×, because methods allocated a scope per call while subs took one from a pool. Now both use the pool. A call-site cache was built and measured neutral, so it was reverted. Object construction, rebuilt per class on every
.new, was levelled with Rakudo on 9 September. - Every
nextin the mainline threw. A sentinel collided with the mainline's frame index 0, so everynextoutside a routine took the exception path, at about 80 µs each: 100,000 of them took 8.1 s.for ^200000 { next if $_ % 2 }went from 4.60 s to 0.03 s, andgiven/when, the same defect, from 4.70 s to 0.10 s. - An argument list stopped being a
std::vector: growth that relocates in one pass, and a free list for small blocks. - A slab for payloads. The allocator was 21.5% of an array-and-hash workload's profile. A slab bought about 4% on regex. It had been predicted at 12%: an allocator micro-benchmark overstates its own effect about threefold.
- An object's attributes stop costing a 4 KB block. Every object's attribute table is a hash, and every hash's first block was 4 KB. A program holding 100,000 objects and 50,000 small hashes went from 531 MB peak to 175 MB. The same change had been tried on 30 August, cost the hash kernel 17%, and was reverted. It landed a month later in a different shape.
A correction worth keeping
On 22 August the benchmarks page said the interpreter had overtaken Rakudo on fib. The Rakudo in that comparison had been running under Rosetta 2, emulated. Re-measured natively, fib was still behind (369.8 ms against 311.4 at v3.7.0). The one-machine re-measurement of every release, two chapters on, exists because of that mistake.