Chapter 4 · July – September
Four ways to run a program
Raku++ was a compiler from its first commit, not only an interpreter. The first commit already had four ways to run a program. Getting the compiled path to agree with the interpreter on every program, then making its binaries small, then letting a running program compile its own hot loops, took the rest of the three months.
rakupp program.raku # interpret
rakupp --bundle program.raku -o program # the interpreter plus the source, as one binary
rakupp --aot program.raku -o program # parsed ahead of time, the tree embedded
rakupp --exe program.raku -o program # transpiled to C++ and compiled natively
--exe possible was something left out. Raku lets a program change its own grammar while it is being parsed: custom operators at parse time, slangs. Raku++ left those out at first. If the parse tree cannot change at run time, it can be turned into C++ at build time. Much later the restraint was lifted on purpose, when slangs arrived, without giving up compilation.Correct by parity
The compiler is not checked against Roast. It is checked against the interpreter: compile a program, run it, run the same program interpreted, and require identical output. The interpreter is the oracle for the compiler.
The corpus for that check was examples/, two dozen small programs: Mandelbrot in ASCII, Life on a torus, a JSON parser written as a grammar, a Brainfuck interpreter, a quine. By 12 July all 23 compiled natively with identical output. The next day's sweep over Roast found that 389 of the 416 fully passing files would transpile.
What made the parameters compile was a calling convention. Every compiled sub takes a uniform ValueList, and binding is emitted per parameter from its signature, so named, optional, default, slurpy and multi parameters all compile rather than falling back.
How much faster native is
| example, 12 July | interpreted | --exe | |
|---|---|---|---|
| life | 746.2 ms | 171.6 ms | 4.3× |
| nqueens | 70.2 ms | 27.9 ms | 2.5× |
| mandel | 106.6 ms | 45.8 ms | 2.3× |
| brainfuck | 7.5 ms | 8.2 ms | 0.9× |
Brainfuck is slower native, and the guide says so: "known, small, and honest". By September --exe ran the benchmark kernels 1.0 to 11.8 times faster than the interpreter. The distance between the two is what the last chapter set out to close from the interpreter's side.
The embedded tree had been losing fields
--aot used to write one builder function per kind of tree node, and any field nobody listed was silently dropped. my Int(Str) $n = "42" compiled into a binary that threw a type error, and compiled programs lost their line numbers. On 1 August the emitter was replaced by a serializer whose reader and writer are the same visitor instantiated twice, so a field cannot be saved and not restored. The same code became the precompilation cache for modules.
Only what the program needs
Until v3.14.0 every --exe binary was the same size whatever the program did: say "Hello" weighed 9.83 MB. --slim split the runtime into five archives and made four features cuttable: Unicode names, collation, properties, and EVAL. A program that turns out to need a cut feature at run time gets X::Feature::NotBuilt rather than a wrong answer. The differential gate compared 270 programs: 241 byte-identical, none different.
say "Hello", compiled. The full binary and the slim one, release by release, as the release notes recorded them.The line climbs after v3.14.0 because the engine kept growing, and a size budget kept count. It was re-pinned four times in six weeks, and once in an emergency: the JIT had pulled the 663,096-byte C++ code generator into every binary, say "Hello" included, for a call that binary could never make.
Compiling while the program runs
On 19 September two opt-in flags let a hot loop stop being interpreted. --jit hands the loop to the C++ code generator and a compiler. --cnp needs no compiler: it copies and patches machine-code stencils the binary already carries, 53 of them. A five-million-iteration integer loop that takes 0.96 s interpreted takes 0.04 s with --jit (warm cache), 0.05 s with --cnp, and 0.03 s as --exe -O. Both flags are marked work in progress. The plan is for --cnp to become the default and for --jit to go.
The JIT had its own lesson. After an afternoon its cache was 91 MB, 89 of it headers. The note it left: an optimisation that spends the user's disk is not free just because it was measured in seconds.