A few months ago the posture inside American companies was simple. Buy the best model money could rent, point it at everything, and let the tokens fly. The Wall Street Journal now calls the replacement posture “thrift-maxxing,” and the description is precise: teams route each task to the cheapest model that can actually finish it, mixing cheaper open models out of China with the offerings from OpenAI and Anthropic. The best is no longer the default. The best is the exception, reserved for the work that earns it, and that single change in buying behavior does more to reprice this industry than any benchmark result has.
The reason it matters so much is that it attacks the assumption underneath the valuations rather than the technology itself. Nothing about the models got worse. What changed is that the buyer learned to ask which tier a given job requires, and once that question gets asked routinely, it never stops being asked.
From Tokenmaxxing to Thrift-Maxxing
The earlier posture was not stupid, and it is worth understanding why it made sense at the time. When nobody knows which tasks a model can handle, buying the top tier for everything is a rational way to purchase certainty: it removes an entire class of failure from the evaluation, and in a period when the technology was being proven internally, removing a failure mode was worth far more than the cost difference. That is what an experimental budget is for.
The problem is that the same behavior becomes indefensible the moment the experiment ends. Once teams have run enough jobs to know that a routine summarization completes fine on a cheap model, the premium is no longer buying certainty. It is buying a habit. A procurement lead who has learned that a two-dollar summarization does not need a twenty-dollar model is a procurement lead who will not pay twenty dollars again, and that lesson does not unlearn itself when the vendor’s next release ships.
Why Procurement Rewrites the Lab Valuations
That one change rattles the math behind the lab valuations in a way that is easy to underweight because it happens quietly and in aggregate. OpenAI and Anthropic are edging toward public markets on a story of endless, undiscriminating demand for the top tier, which is a story about volume at a given price. The thrift-maxxing buyer does not reject the product. They simply decline to buy the expensive version of it for the majority of the work, and the majority of the work is where the volume lives.
The IPO slides still show the curve climbing forever, and the curve may well climb. The procurement lead is quietly rewriting the assumptions underneath it anyway, one routing rule at a time, and routing rules do not appear in anyone’s forecast until they show up as a revenue mix that nobody modeled. That is the gap worth watching: not a demand collapse, which would be visible and dramatic, but a slow migration of the median task down the price ladder while the headline number still grows.
None of this is the boom failing. It is the boom growing up, and the history here is not subtle. Railways got laid in a frenzy, and then the accountants arrived to make the lines pay. Cloud went from “move everything, cost be damned” to FinOps teams clawing the bill back line by line. The moment buyers start optimizing is the moment a technology stops being a bet and becomes a tool, which is the transition every serious technology has to survive. Discipline in the buyer is a sign of health rather than decay, and it is usually mistaken for the opposite by the people selling.
Reliability Is the Moat, Not Capability
The same discipline shows up in the questions nobody has cleanly answered yet. The New York Times experiments on whether AI could stand in for an office worker produced a result that read less like a verdict than a confession: it can do a great deal, and it still needs someone watching. That is not a failing grade. It is a description of a tool that shifts where the labor goes rather than removing it, which is exactly the kind of finding a procurement team can act on and a marketing team cannot use.
Set that against the parallel observation that the fundamental challenge of software development has moved, and the two together describe the whole market. The puzzle is no longer whether a model can write the code. It is whether you can ship something you would actually trust to run in production on a Friday. Capability has commoditized at the bottom of the stack, fast and thoroughly. Reliability has not commoditized at all, and there is no sign it is about to.
So the moat was never raw capability, which was always the most copyable thing in the building. It is the unglamorous property: does the system behave the same way twice, and who answers the phone when it does not. A lab that earns its premium on the tasks that genuinely require the top tier will keep that premium, because those tasks are real and the buyer can identify them now. A lab that was selling “always buy our best” just lost its easiest customer. Thrift does not kill the leaders. It separates the ones charging for a brand from the ones charging for a result, and that separation was always coming.
There is a pattern running underneath all of this that the room tends to laugh at before it concedes. The thing that turns out to be true usually gets dismissed first, and the dismissal comes loudest from the people with the most invested in the old story. Cheap open models would never matter to a serious enterprise. The office would either vanish by spring or the entire category was hype. Each claim got its laugh in its moment, and each one is now something procurement teams act on directly. The laughter was the tell. When the obvious gets ridiculed, it is generally because it is about to be proven right by the people doing the work rather than the people selling the dream.
Accountability Moved Downstream
The labs are not finished, and nothing here suggests they are. The ones with taste and execution, the ones willing to cut ruthlessly and ship something that holds up under load, will do fine on the far side of this. What ends is the free ride of unexamined spend, which was never a business model so much as a phase. A market that makes its buyers choose is a market that has finally started to respect them, and respect is more durable than enthusiasm.
The part nobody put in the slide deck is what came attached to the choice. The buyers chose, and with the choice came the consequence, and the consequence no longer stops at the lab’s door. The company that routed the cheap model owns what that model produced. The team that shipped the AI-written code owns the bug at two in the morning, and owns the explanation afterward. Accountability moved downstream to the people actually using the thing, which is the quiet structural shift underneath the cost story and the more important half of it.
A technology stops being a story the moment somebody has to make it pay. The AI boom just got its first real invoice, and the bill and the blame now land in the same inbox.

Leave a Reply