llama.cpp Put Precision Policy Inside the Model. That Matters More Than W4A4
A llama.cpp change merged on September 25 looks like a narrow Blackwell optimization. It is actually a useful preview of how local inference should work. The runtime can now read precision policy from model metadata and choose a different computation path for specific tensors. That sounds like plumbing. I think it moves an important decision out of deployment folklore and into the model artifact itself.
The immediate problem involves NVFP4 models on NVIDIA Blackwell. Their weights use four bit precision. Blackwell can also process activations at four bit precision, creating the fast W4A4 path. But not every layer tolerates that treatment equally well. Some layers were prepared with an expectation closer to W4A16, and forcing their activations down to four bits can hurt output quality. The fastest native path is not automatically the correct path for every tensor.
One model can need more than one precision
The merged change adds metadata that identifies the affected layers. llama.cpp reads that metadata and routes those layers through a W4A8 path while leaving other layers eligible for the native W4A4 path. The pull request includes support for both dense matrix operations and mixture of experts operations. It was tested with an NVIDIA Qwen3.6 35B A3B NVFP4 model.
That is the technical change. The architectural change is more interesting. Precision is no longer treated only as one global choice made when a server starts. The model can carry information about where aggressive activation quantization is acceptable and where it is not. The runtime becomes responsible for enforcing that policy on the hardware that can use it.
I like this direction because quantization has been sold as a label for too long. People see four bit weights and assume they understand the speed, memory use, and quality profile. They do not. Two files with the same weight precision can behave differently because activation precision, sensitive layers, kernels, and hardware paths all matter. A single label hides the decisions that actually determine whether the model is useful.
The fastest kernel can be the wrong kernel
The llama.cpp pull request reports that using W4A8 for the marked layers improved perplexity on its test model compared with the native W4A4 baseline. It also reports lower prompt processing throughput in that specific comparison, while text generation speed was much closer. Those are project supplied measurements, not results from my RTX PRO 6000, so I would not generalize the numbers. The direction is enough to make the point: the quality and speed tradeoff can differ by layer and by phase of inference.
That distinction matters for agents. A background agent does not get much value from shaving a small amount of time off generation if lower precision causes one more bad tool choice or one more malformed structured response. Interactive work may value prompt processing latency more. The right precision policy depends on the workload, but the runtime needs enough model specific information to make a sensible choice before any workload tuning begins.
The old model was simple. Pick a quantization, load it, and assume the runtime should use the most aggressive supported kernels. The new model is more honest. A model is a collection of tensors with different tolerances, executed through a stack with different hardware capabilities. Global flags are too blunt to express that cleanly.
Metadata is becoming executable infrastructure
GGUF metadata used to feel like packaging information. Architecture, tokenizer settings, context length, and tensor descriptions helped the runtime understand what it had loaded. Precision policy makes the metadata more operational. It can influence which computation path runs on the GPU.
That makes the model file closer to a deployable artifact and less like a bag of compressed weights. The conversion pipeline records an intent. The runtime interprets it. The backend maps it to an available kernel. If another runtime ignores the same metadata, the identical weights may produce a different quality and performance profile.
This is where model portability gets harder and better at the same time. Better, because more knowledge can travel with the model instead of living in a README or a launch script. Harder, because supporting a file format is not the same as supporting every policy encoded in it. A runtime may load a model successfully and still miss the execution behavior its creator expected.
For a founder building model agnostic systems, this is a reminder that compatibility has layers. An OpenAI compatible endpoint makes APIs portable. A shared model format makes artifacts portable. Neither guarantees equivalent execution. The runtime and hardware still decide how those artifacts become computation.
Local inference is becoming a policy engine
The interesting local AI competition is moving above raw kernel support. Runtimes already decide memory placement, cache layout, batching, speculative decoding, and device assignment. Precision policy belongs on that list. The mature runtime will not merely expose every possible flag. It will combine model metadata, hardware capability, and operator preferences without making the operator memorize which layers fail under which path.
There is a limit to how much should be automatic. Metadata can be wrong. Conversion tools can lag behind models. A runtime update can change behavior even when the model filename stays the same. The new llama.cpp build documentation keeps explicit choices available: q4 can force the native W4A4 path, q8 can force W4A8 for every layer, and auto follows the per tensor metadata. That escape hatch is important because automatic policy should be inspectable and overridable.
I would not call this proof that mixed precision has been solved. It is one implementation for a specific NVFP4 path on Blackwell, tested on one named model in the pull request. But it is the right abstraction. Put model specific precision knowledge in the artifact, let the runtime map it to hardware, and preserve an override for measurement and debugging.
The test I care about now is straightforward: run the same NVFP4 model with auto, forced q4, and forced q8 on a Blackwell system, then compare task quality, prompt processing latency, and generation speed on the actual agent workload. If auto keeps most of the speed while recovering the failures that appear under forced q4, precision metadata has graduated from a format detail into part of the deployment contract.