Chapter 3 · July – September
Exact numbers, whole graphemes
Raku promises that 0.1 + 0.2 == 0.3 is True, because its decimals are exact rationals. It also counts a string's characters as a reader sees them, not as bytes or code points. Both promises are invisible until they are broken, and both were kept from the first commit. July was spent making them complete, and the rest of the time making them fast.
Whole graphemes
- Unicode is the strongest area from the start: 40 files in full, 95% of the synopsis's assertions.
- Unicode number literals take the literals section to 100%.
- Unicode 17 grapheme clustering (UAX #29) and the four normalization forms: S15 at 99.9%, and
emoji-test.twith its 3,825 tests passes. Then UCA collation from DUCET 17.0, and all 8,271 conformance tests pass. That took the whole suite from 76.2% to 80.6% of its tests. - Strings are normalized when they are built, as Raku requires. That made
~=quadratic for a day, until strcat went from 360 ms back to about 20. - 100% Unicode. 91,752 of 91,752 assertions in S15. The one file not counted,
concat-stable.t, is a timeout, not a failure.
Exact numbers
Underneath every arithmetic operation is a hand-written big integer, base 109 so that printing a number needs no conversion, with a fast path while it fits in 64 bits. A Rat is two of them. Raku also says when exactness ends: a Rat whose reduced denominator no longer fits in 64 bits becomes a Num, and FatRat never does. That rule arrived on 11 July with rat.t (869 tests) and complex.t (557) passing in full.
The Mandelbrot
examples/mandel.raku descends from the ASCII Mandelbrot that shipped with Parrot two decades ago. Its decimal literals are Rats, so the render is exact rational arithmetic, and it ran slower than Rakudo. One day's work on 11 July followed the profile:
- 6.83 s → 3.62 s: a 128-bit fast path for small Rats, and Rat literals cached on the tree node. A bare
10.0in a hot loop had been allocating two big integers and running a GCD every time it was evaluated, about 10 µs each. - 3.6 s → 0.32 s (0.21 s on arm64): the fast path's numerator and denominator were already coprime, and the constructor reduced them again anyway. GCD on already-reduced 19-digit parts is its worst case: about 90 allocating steps per operation.
The compiled binary took 0.12 s. The same maths, the same fractal, drawn in a blink.
Fast on other people's workloads
- A 64-bit fast path through GCD and division. GCD went from 15,761 ns to 183, and a Rat-heavy sum loop from 717 ms to 72. The old division had done a binary search over [0, 109) for every limb.
- Knuth's algorithm D. A
Math::NumberTheorytest never finished: not slow, never. Reducing a 1,437-digit number over an 812-digit one took 865 ms, and takes 7 now. - An issue report timed a graph library at 4× Rakudo. 128-bit fast paths took
%from 4.1× Rakudo's time to 1.4×, anddivfrom 6.0× to 2.1×. The same benchmark exposed four silent wrong answers: a graph's diameter came out 399 instead of 38. The report put it this way: it had looked faster only because it was not doing the work. - Eight independent carry chains in multiplication. The big-integer kernel went from 13.0 to 7.4 ms. One chain costs about five cycles a limb where the instruction count says one and a half; four or six chains perform within 3% of eight, and sixteen is too many.
- The first four limbs are stored inline, so any integer below 1036 never allocates. The Rat kernel went from 2.00 million to 1.20 million allocations.
Division, according to the VM
On 13 September a module that calls the VM's own opcodes showed how they divide. nqp::div_i floors (-7 div 2 is −4), nqp::mod_i truncates (−1), and mod_I floors again. Raku++ had floored mod_i as well, and answered 0 on a zero divisor. No Roast file calls an nqp opcode, so nothing had noticed. All three now do what MoarVM does.