Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It
The cache policy that won in simulation is patched into a live vLLM and measured on real hardware. The simulated win vanished at first; after fixing where and what to measure, tail latency (p99 TTFT) came out 13-16% lower under memory pressure. The prototype's costs and the conditions where it does not help are reported as measured.
Read →