vLLM 0.28 Turns KV Cache Into a Storage Hierarchy. That Changes How I Think About Local AI Memory
vLLM released version 0.28 this morning with 584 commits from 270 contributors. The headline items cover Kimi K3, DeepSeek V4, speculative decoding, a newer model runner, and a Rust front end. The feature I care about most is quieter: tiered KV cache offloading now includes disk. That sounds like