Notes
What a measurement taught me
Short notes on one thing each. The tables, raw data and limits are in the repositories they link to.
-
A simulation can't see a cache
My gateway's failure tests all passed. A routing rule I had added still cut prefix-cache hits from 59.1% to 24.4% on real engines.
-
The wrong answers that raise no error
A fine-tuned 0.5B model writes the right SQL 82.6% of the time. What worries me is the 16.4% that runs cleanly and returns the wrong rows.
-
Same weights, different dtype
Only 27.6% of a fine-tuned model's bytes match its base, and the encoding mattered more than the chunker.
-
Average chunk size isn't what an edit costs
An edit is more likely to land in a big chunk, so the spread of chunk sizes matters as much as their average.