Yuchen He 中文

Notes

What a measurement taught me

Short notes on one thing each. The tables, raw data and limits are in the repositories they link to.

  1. A simulation can't see a cache

    My gateway's failure tests all passed. A routing rule I had added still cut prefix-cache hits from 59.1% to 24.4% on real engines.

  2. The wrong answers that raise no error

    A fine-tuned 0.5B model writes the right SQL 82.6% of the time. What worries me is the 16.4% that runs cleanly and returns the wrong rows.

  3. Same weights, different dtype

    Only 27.6% of a fine-tuned model's bytes match its base, and the encoding mattered more than the chunker.

  4. Average chunk size isn't what an edit costs

    An edit is more likely to land in a big chunk, so the spread of chunk sizes matters as much as their average.