Yuchen He 中文

Notes

A simulation can't see a cache

My gateway's failure tests all passed. A routing rule I had added still cut prefix-cache hits from 59.1% to 24.4% on real engines.

My gateway’s failure benchmarks all passed. They use simulated workers, and a simulated worker has no KV cache, so they could not tell me that a routing rule I had added was bad.

The rule: when a request’s preferred node is full, spill to the next-ranked one. It passed the tests. Then I ran the policies against three real llama.cpp servers: 120 requests over six long shared prompts, caches erased before every run.

Originalhash % N, queues when full59.1%
Spill when fullthe upgrade's first rule24.4%
Placement + waitthe fix61.5%
Round-robinno gateway22.0%
Prompt tokens read from the KV cache, by routing policy. Three llama.cpp workers (Qwen2.5-0.5B, two slots each) on an Apple M1; 120 requests over six shared long prompts; cache capacity limited to the slots; caches erased before every run; median of 3 runs. Source: radixgates, benchmarks/results/real_engine_limited_cache/.

Spilling scattered each prefix over several nodes and evicted it from all of them. The original gateway, which just queues, kept 59.1% of its prompt tokens in cache, but sent three quarters of the prefill work to one worker. Placement memory plus a bounded affinity wait matched the original’s hit rate (61.5% against 59.1%, indistinguishable over three runs) and spread the work: the busiest worker’s share went from 73% to 43%. That is parity on cache hits, not a win.

What I took from it: a simulation only tests what you put in it. The failure I cared about lived in the piece I had left out, so after any simulated result the next step is the smallest real system that contains that piece. And the fix is off by default, because with llama.cpp’s own host-memory cache every policy reaches 95 to 98% and routing barely matters.

The per-policy table, the raw runs and the limits (one machine, three runs, a 0.5B model) are in the radixgates README.