Chapter 15 · 30 September

The tree-walker, 20% faster in a day

From v5.1.0 (4c89cbae) to b81c74fd: thirteen engine commits working through the plan's tasks 2–9. The gap between interpreting a program and running its --exe build shrank by a third to a half on the kernels where it was widest, and Roast stayed at every file passing.

−20%mean time over the 19 perf-guard kernels, vs v5.1.0
4.3×fib behind --exe, from 7.6×
−58%streq, the program with the widest gap
175 MBpeak memory holding 100k objects, from 531 MB
1,424 / 1,424Roast files passing, after every commit

How far behind native code the interpreter runs

Interpreted time divided by the time of the same program compiled with --exe. 1× would be native speed.

v5.1.0now
Table: the gap in milliseconds

The benchmark programs, against v4.0.1

Each program's time as a percentage of v4.0.1's (the dashed line). Left of the line is faster. The line from the orange dot to the blue one is today's work. Sorted by today's time.

v5.1.0nowv4.0.1 = 100%
Table: milliseconds, best of 5

The perf-guard kernels, v5.1.0 to now

Change in time on each of the 19 release-gate kernels. junctionwide runs for 18 ms and is the noisiest in the set.

Table: milliseconds, best of 5

What landed

Each figure is that commit's own interleaved A/B against the commit before it, best of 5 or 7.

commitchangemeasured

Tried and dropped

  • A small front function ahead of evalBinary: fib +4–5%, two calls where there was one.
  • Moving evalBinary's metaop or DateTime block out: its frame shrank by 0 and 48 bytes. It has no dominant arm.
  • A second compiled handler for fib($n-1) + fib($n-2): +1.1%, the code it jumped to has as large a frame as the one it skipped.
  • An inline Int switch in the fast shape, a one-byte metaop gate, a first-byte reject in isNativeScalarName: no measurable change.

How it was measured

  • Local arm64 Release builds of v4.0.1, v5.1.0 and HEAD on one machine, run in turn (A, B, C, A, B, C…), best of 5.
  • Another process kept the load average at 4–8 during the final run; interleaving gives every build the same conditions, so the ratios hold.
  • Every commit passed Roast and the source size budget. The corpus diff and the module-battery suites ran clean on all of it through 8fd1e905.
  • Still slower than v4.0.1: textsplit +8% and multiwhere +5%. Both regressed gradually over the Roast bursts; about half of it is won back.