Kimi K3 Shows Why Speculative Decoding Is a Memory Decision

Share

llama.cpp merged recurrent state rollback support for Kimi K3 on September 8. The code change sounds narrow, but the measurements attached to it expose a broader mistake in how people evaluate speculative decoding. We talk about speculation as a speed feature. On large recurrent models, it is also a memory allocation and concurrency decision.

The merged pull request reports tests on Kimi K3 using eight RTX PRO 6000 Blackwell GPUs, four server slots, and seven draft tokens. The results are not a clean product benchmark, and the author explicitly documents differences in model placement. They are still useful because they show the trade in unusually concrete terms. The four slot recurrent cache grew from 1772.44 MiB to 14179.50 MiB when seven rollback positions were stored.

That is the part local AI builders should notice. Speculative decoding can make one request faster while making the serving system worse at the workload you actually bought it to run.

Recurrent models change the rollback problem

A conventional transformer usually handles generation state through a key value cache. Recurrent and hybrid architectures carry additional state forward as tokens are processed. Kimi K3 stores KDA state plus convolution windows. When a draft model proposes several tokens, the target model evaluates them. If it rejects part of the proposal, the runtime must restore the target to the correct earlier state before continuing.

Before this change, Kimi K3 had recurrent rollback disabled in llama.cpp. Simply turning it on would not have been safe because the runtime stored only the final KDA state and convolution windows. Rejecting draft tokens could therefore restore snapshot groups that had never been written. The merged implementation saves the convolution windows for every rollback position, saves KDA state snapshots, and adds Kimi K3 to the supported architecture list.

This is a reliability change before it is a performance change. A faster token stream is worthless if rejection leaves hidden recurrent state inconsistent with the visible sequence. The pull request added tests for checkpoint restoration, replay across multiple sequences, and sequence isolation. That testing scope is a good model for anyone deploying speculative decoding in an agent system. Validate the state transition, not only the speed counter.

The memory cost scales with ambition

The pull request states that snapshot storage multiplies recurrent state memory by one plus the number of rollback positions. Seven positions therefore mean eight copies of the relevant state. In the recorded four slot configuration, that took the cache from about 1.77 GiB to about 14.18 GiB.

More draft tokens create more opportunity to skip target model work, but they also require more state that can be restored after a rejection. More server slots multiply the serving footprint again. This is why the right setting cannot be copied from a single user demo. A workstation serving one interactive session has a different optimum from a box running several background agents.

VRAM is not merely where model weights live. It is the operating budget for context, recurrent state, draft models, server slots, and runtime overhead. A configuration that fits the weights with a few gigabytes left is not necessarily ready for speculation. The feature may consume the margin that would otherwise support longer context or another concurrent agent.

Faster at one request, slower at four

The concurrency results make the decision clearer. In the reported random input benchmark, the combined draft model and rollback setup reduced time per output token from 77.7 milliseconds to 49.2 milliseconds at concurrency one. At concurrency four, time per output token increased from 134.5 milliseconds without speculation to 220.6 milliseconds with the combined setup.

These rows do not isolate rollback from the draft model, and the pull request warns that host checkpoint results were unavailable for these tests. Treat them as a workload signal, not a universal ratio. The signal is still hard to miss. Speculation improved the isolated request and lost badly when all four slots were active.

That pattern matters for autonomous agents. A coding assistant used by one person may benefit from lower latency on each generation. A background agent host usually cares about completed tasks across a time window. If speculation reserves enough memory or compute to reduce useful concurrency, the faster stream can lower total system throughput.

Benchmark completed work per machine

The default speculative decoding benchmark should not be one prompt and one generated token rate. Start with the actual number of simultaneous agents. Keep the model, quantization, context distribution, output length, and prompt caching behavior fixed. Record total VRAM, accepted draft tokens, time to first token, time per output token, requests completed per hour, and failures. Then repeat without speculation.

For agent workloads, add task completion time. Agents alternate generation with tool calls, retrieval, and prompt growth. A decode optimization that looks large in isolation may barely move the full task. Conversely, a configuration that preserves another server slot may complete more work even if each visible response arrives more slowly.

The merged Kimi K3 change also shows why architecture support belongs in model selection. An open weight model is not fully usable just because a runtime can load it. Correct rollback, batching, quantization, and cache behavior determine whether it can serve the intended workload. Those capabilities often arrive after the model release, one careful pull request at a time.

Buy memory for the serving plan

My conclusion is not that speculative decoding is bad. It is that the feature spends resources to buy latency, and the exchange rate changes with concurrency. On a high memory workstation, the right answer may be to use it for interactive work and disable it for a pool of background agents. On a smaller card, the snapshots may rule it out before performance testing begins.

The practical test is simple. Run the same agent queue at concurrency one and at your real operating concurrency, with and without speculation. If completed tasks per hour rise without pushing the system into an unsafe memory margin, keep it. If only the single request token counter improves, spend the VRAM on another slot instead.