Claude Sonnet 5.5 Makes Model Routing More Valuable Than Model Loyalty
Anthropic released Claude Sonnet 5.5 on September 28. The company calls it a clear upgrade over Sonnet 5, says it runs 30 percent faster, and says it costs up to 30 percent less for most work. That is a meaningful product improvement. It is also another reason I think model loyalty is becoming a bad architectural choice. The faster model cycle gets, the less sense it makes to design a workflow around the assumption that today's winner will still be the right model six months from now.
The easy reaction is to ask whether Sonnet 5.5 beats the other frontier models. That question matters for a benchmark table and for a few demanding tasks. It matters much less for the architecture of an actual product. A production system needs reliable outputs, predictable latency, acceptable cost, tool support, and a fallback when the preferred endpoint fails. No single model owns all of those properties. More importantly, the answer changes too quickly for model choice to become a permanent design decision.
A better model is not a better dependency
Sonnet 5.5 can be better than Sonnet 5 without making Anthropic a better place to store workflow logic. Those are separate judgments. Model capability belongs in an evaluation. Provider dependence belongs in an architecture review. Mixing them is how a temporary quality lead turns into a permanent switching cost.
The lock in usually does not begin with the application interface itself. Most providers expose familiar chat and response patterns. It begins in the details around the model: proprietary tool definitions, provider specific prompt behavior, cache assumptions, tracing formats, retry logic, and evaluation data that never gets separated from one vendor's dashboard. Each shortcut looks reasonable in isolation. Together they make the next model comparison expensive enough that nobody wants to run it.
That is why a release like Sonnet 5.5 should increase the value of a routing layer rather than reduce it. Faster and cheaper inference expands the number of tasks where the model may win. It does not remove the need to compare it against an open model running locally, a lower cost hosted model, or another frontier endpoint. A router preserves that comparison as an operating capability instead of a migration project.
The frontier premium is now task specific
There was a period when paying for the strongest closed model was an easy default. The quality gap covered a lot of architectural laziness. That gap is no longer uniform. Open weight models can handle a growing share of extraction, classification, routine coding, document work, and tool use. Local inference adds privacy, availability, and fixed capacity. Frontier systems still earn their premium on hard reasoning, difficult coding, and tasks where a small quality difference changes the outcome. The premium now belongs to particular tasks, not to an entire application.
This distinction changes the economics. A 30 percent improvement in speed is valuable when a user is waiting on an interactive answer. It may be nearly irrelevant for a background agent that can finish overnight. A higher price can be rational for one difficult planning step while being wasteful for fifty mechanical tool calls that follow. One model can plan, another can execute, and a third can check the result. The model boundary should follow the work rather than the logo on the account.
I run local and hosted models because they solve different operational problems. The local systems provide capacity I control and let background work continue without turning every loop into a metered event. Hosted frontier models are useful when the capability premium is real. The important part is not declaring one side the winner. It is keeping the workflow portable enough that each step can move when the quality, latency, or price curve changes.
Release velocity punishes hard coded choices
Sonnet 5.5 arrived as a direct improvement only months after Sonnet 5. Open model families are moving just as quickly, while inference engines such as vLLM, SGLang, llama.cpp, and MLX keep changing what existing weights can do on available hardware. A model comparison is therefore a dated measurement, not a durable fact. The winner can change because a new model ships, because a serving engine improves, or because a quantization finally fits the hardware already in the rack.
This is why I care more about repeatable evaluations than launch benchmarks. Published benchmarks describe a model in a controlled test. A workflow evaluation measures whether the model produces the structure, tool calls, judgment, and failure rate the product actually needs. If that evaluation can run against several endpoints, Sonnet 5.5 is easy to adopt where it wins. If the evaluation is inseparable from one provider, even a cheaper model can increase the real cost of the system.
Routing does add complexity. Prompts do not behave identically across models. Tool calling formats differ. More endpoints create more failure modes, and local inference requires operational attention. A router that blindly sends every request to the cheapest model is not architecture. It is a billing trick. The useful version has explicit task classes, measured quality floors, clear fallbacks, and records that show why a request went where it did.
The durable asset is the workflow
Models are improving too quickly to be the stable center of an AI product. The durable pieces are the workflow, proprietary context, evaluations, permissions, user trust, and distribution. Those survive a model swap. Sonnet 5.5 strengthens this argument because it is good news for customers and bad news for anyone who treated Sonnet 5 as a permanent platform. The next upgrade will repeat the lesson.
My recommendation is to treat Sonnet 5.5 as a candidate, not a commitment. Run the same real task through the current production model, Sonnet 5.5, and at least one open weight alternative. Measure quality, latency, failure recovery, and total cost at the workflow level. Then switch the route where the evidence supports it. If adopting a better model requires redesigning the product, the model was never the main problem. The architecture was.