A Three Line llama.cpp Fix Explains Why Two GPUs Are Not One Bigger GPU
A llama.cpp change merged this morning removed one guard, added one explicit device selection, and enabled an existing CUDA graph optimization on systems using more than one GPU. The code change was three lines added and one removed. The result was not a universal speedup. It improved token generation by about 3.4 to 3.7 percent on one mixture of experts model, while a dense model stayed effectively flat.
That small result explains a large hardware mistake. Two GPUs are not one bigger GPU. They are two devices connected through a software path, and every model architecture can use that path differently. Buying the second card is easy. Proving that your actual workload benefits from it is the engineering work.
What the patch actually changed
The contributor behind llama.cpp pull request 28198 found that systems with multiple GPUs were skipping the graph optimization for concurrent CUDA streams. The computation graph was already divided by device, and the optimization ran separately on each split. The blanket guard against multiple devices was therefore unnecessary.
There was still a device ownership problem. A CUDA event belonged to whichever GPU was current when the event was created. If the optimization ran while GPU zero was current, an event intended for the second GPU could land on the first. The patch removed the device count restriction and explicitly selected the correct device before creating the stream context. The feature remains opt in through the CUDA graph optimization setting, so default behavior did not change.
This is the kind of implementation detail hidden behind a product page that says multi GPU support. A model may load across two cards and produce correct output while leaving a useful execution path disabled. Capacity can work before concurrency works. Correctness can work before performance works.
The benchmark distinction matters
The contributor tested an RTX 5070 Ti with 16 GB and an RTX 3080 Ti with 12 GB. The second card used a slower chipset slot. There was no NVLink and no direct peer access. About 6.1 GB of the model buffer landed on one card and about 7.4 GB on the other. With the optimization enabled, the logs showed 50 concurrent stream launches instead of zero.
The correctness checks were stronger than a quick prompt comparison. Greedy generation produced byte identical output in both configurations. Perplexity matched for every chunk in an eight chunk WikiText test. Those checks matter because a faster answer is worthless if an execution change quietly alters model output.
Then the two model architectures diverged. On a quantized Gemma 4 mixture of experts model, prompt processing changed by only 0.2 percent, but token generation improved by roughly 3.4 to 3.7 percent across three generation lengths. On a quantized Qwen 3.8 dense model, prompt processing moved by 0.1 percent and generation moved slightly backward, by 0.1 to 0.2 percent. These are the contributor's measurements on that specific mixed GPU machine, not results from my hardware and not a promise for another configuration.
The important result is not 3.7 percent. It is the split between architectures. The same runtime change, on the same machine, was useful for one model and irrelevant for another. A generic multi GPU benchmark would hide the decision a buyer actually needs to make.
Capacity and speed are separate purchases
A second GPU can solve a capacity problem even when it does not improve generation speed. If a model does not fit on one card, splitting it may be the only practical route. That is a valid purchase. But it should be described honestly as buying access to a larger model, not automatically buying lower latency or greater throughput.
Performance depends on what crosses the device boundary, how layers are divided, whether the workload is dense or sparse, the interconnect, slot bandwidth, peer access, runtime scheduling, and which optimizations support the split path. A fast card in a constrained slot can still add valuable memory. It can also create a communication boundary that erases expected compute gains.
This is why I do not like comparing local AI machines by adding their advertised compute or memory bandwidth. Those totals describe components. They do not describe the path a token takes. The execution path is the product.
The same rule applies beyond two cards in one workstation. Two compact systems linked over a network provide aggregate memory and independent capacity, but they do not become a single accelerator because a framework can see both. Running separate agent workers may use two machines efficiently. Splitting one latency sensitive model across them may expose communication costs. Architecture decides whether parallel hardware should cooperate on one request or handle different requests.
What founders should measure before buying card two
Start by naming the constraint. If the model does not fit, measure whether the split makes it usable and stable. If latency is the problem, measure prompt processing and generation separately. If throughput is the problem, test concurrent requests rather than one chat. If the target is autonomous agents, measure completed jobs per hour, queue delay, recovery from failures, and useful work per dollar.
Then repeat the test with at least one dense model and one mixture of experts model you would actually deploy. Record the exact quantization, context length, batch size, tensor split, slot topology, runtime commit, and optimization settings. Verify output quality or perplexity before accepting a faster number. A benchmark without the execution configuration is not evidence that another buyer can use.
I would also compare the split configuration with independent serving. Put the full model across both devices for one test, then run a smaller model per device and route requests between them. The split setup may win on model capability. Independent replicas may win on throughput, isolation, and failure handling. For background agents, the second answer can be more valuable than a small gain in single request speed.
The llama.cpp patch is good engineering precisely because its claim is narrow. It fixes the device context, verifies correctness, and shows where performance moved and where it did not. Hardware buyers should use the same discipline. Before buying a second GPU for speed, reproduce your workload on borrowed or rented hardware and require the split setup to beat independent serving on the metric your system actually needs. If it only makes a larger model fit, call that the win.