The End of Model Management
When top-tier AI turns into a commodity, the edge is no longer the model. It is who steers it.
When top-tier AI turns into a commodity, the edge is no longer the model. It is who steers it.
Two dollars. That is what a million tokens of input now cost in Anthropic’s new Claude Sonnet 5, ten dollars for the output, on introductory pricing through the end of August. The flagship, Opus 4.8, costs five and twenty-five. So you get near-Opus capability for roughly forty percent of the price.
The obvious headline is that AI got cheaper again. That reading is correct, and it is the least interesting one available. Because when capability that yesterday lived only in the expensive flagship becomes an affordable commodity overnight, the first thing that changes is not what is possible. It is who holds the advantage. And that advantage moves to exactly the place most companies are not looking.
What actually happened this week
Sonnet 5 is not a benchmark show. On the one coding measure Anthropic reports directly, the model scores 63.2 percent, against 69.2 percent for the larger Opus 4.8. The gap to the top has narrowed, but it has not closed. Anyone waiting for a new record will be disappointed.
The real move sits in the price tag. TechCrunch frames Sonnet 5 as the cheaper way to run agents, VentureBeat reads it as a steep discount on Anthropic’s own flagship, in the middle of a race toward an IPO. On output price, Sonnet 5 lands at a third of OpenAI’s GPT-5.5, which charges thirty dollars per million tokens. That is not a technical detail. It is a declaration that agentic capability is now the baseline expectation at every price tier, and that the competition has shifted to who can deliver it most cheaply and most reliably.
This is the beginning of Work after AI. Not the moment the machine can do everything, but the moment good machines become so cheap that owning one is no longer an edge.
The token price is the wrong number
Here is the part almost no one says out loud. The price per token is no longer a reliable metric. Sonnet 5 uses a new tokenizer that maps the same work onto one to 1.35 times as many tokens. Anthropic set the introductory price, in its own words, to be roughly cost-neutral. The price per token fell, and the tokens per task rose, and the two roughly cancel.
Read that twice, because it breaks the entire way companies have bought AI so far. Two models with an identical token price can cost completely different amounts to finish the same job. A model that answers the same question in more steps and more tokens is the more expensive one, despite the same price list. And the more autonomy you hand a digital colleague, the higher the effort level, the more tokens it burns, not fewer.
So choosing a model from the price list means deciding on the wrong basis. The only figure that matters is cost per outcome. And it appears on no datasheet. It exists only once a concrete task runs through a concrete model at a concrete setting. That is the difference between managing a model and steering intelligence.
The bottleneck moves from the model to the steering
As long as there was one clearly best model, the job was simple. You took the best one. That era ends this week. Between Sonnet 5 and Opus 4.8 you can tune the balance of cost and performance through the effort level. Below them sit Gemini 3.5 Flash and open models like DeepSeek, whose output price runs under a dollar per million tokens, a full order of magnitude beneath Sonnet 5. Above them, the Opus and GPT ceiling at twenty-five and thirty dollars.
Between that floor and that ceiling lies a factor of thirty in price. Open source sets the floor, the frontier sets the ceiling. The question of which single model is best becomes the wrong question. The right one is which intelligence for which task, at what cost, and at what risk. Run routine work on the expensive flagship and you burn money. Hand the delicate judgment call to the cheapest open model and you burn trust.
The market already feels this without naming it. One survey of enterprise teams puts the share of companies actively managing their AI costs at 98 percent for 2026, up from 63 percent in 2025 and 31 percent in 2024. In the same breath, those teams say they can see their spend rising but not who is driving it or what value it creates. That is the exact picture of a bottleneck that has moved. The model is no longer scarce. What is scarce is the ability to steer a whole portfolio of intelligences by task, cost, and risk, and to make that steering accountable.
What this means for owners and capital
For owners and boards, this is the genuinely uncomfortable point. Access to top-tier AI was a differentiator for a while. This week it stops being one. When near-Opus capability is available to everyone at forty percent of the price, owning the model is as much of an edge as owning a power line. Andreessen Horowitz finds that enterprise CIOs expect their generative-AI budgets to grow by roughly seventy-five percent in the coming year. That capital is flowing into a layer that is turning into a commodity.
So value moves up, into the layer above the model. Into the orchestration that decides which task uses which intelligence. Into the context that makes a model useful in the first place. Into the governance that proves the right decision was made for the right reason. It is the parallel to the cloud, whose compute became cheap and whose real discipline afterward was cost management. The model zoo brings its own discipline. A board that looks for its edge in having licensed the most expensive model is confusing a higher bill with a stronger position. It is the same error as mistaking a leaner balance sheet for a better one.
Anyone valuing a company in this cycle should not ask which AI it uses. They should ask who there decides which intelligence does which task, and whether that company even knows its cost per outcome. The answer separates the firms that own AI from the firms that command it. The distance between the two will not be closed by one more model swap.
The new core competence already has a name
That names the shift. Model management, the picking of a model, was the competence of the last three years. Intelligence management, the steering of a portfolio of intelligences by cost, risk, and task, is the competence of the next. The good news for Europe is that this competence plays to its strengths. Documented processes, cost discipline, and governance were long treated as a brake. In a world where capability becomes a commodity and steering it becomes the edge, they turn into a differentiator. Whoever steers intelligence systematically, and can prove it, builds something a competitor cannot simply buy off the shelf.
From here the models get cheaper and better, week after week. The edge no longer lies in owning the best one, but in steering many of them well. Whoever waits will find, a year or two from now, that the decisive competence was available all along, and that a competitor was practicing it while they were still debating the next model. And they will not have seen it coming.
Gerhard Kürner is CEO of 506.ai, the European platform for Service-as-a-Software and agentic engineering.





