The Inference Bill: Why AI's Best Customers Lose the Most Money

Every other software business on earth follows one rule: your heaviest users are your most profitable. Serving one more user of a database or a spreadsheet costs almost nothing, so the more they use it, the wider the margin. AI broke that rule. The best customer of a frontier lab is often its least profitable one, and the reason is a single line item that never appears on the marketing page: the inference bill.


Two costs, and only one of them recurs

There are two bills in the model business, and they behave nothing alike. Training is the cost of building the model once: a long, brutally expensive run that happens before anyone uses anything. It is a capital expense, paid up front, amortized over the model's life. Inference is what happens every time someone asks the model a question. The weights get loaded, the tokens get generated, the electricity gets spent, and it happens again for the next query, and the next, forever.

Training is the culinary school. Inference is the stove firing up on every single order. And it is the inference bill, not the training bill, that scales with success. OpenAI told investors its inference expense rose roughly fourfold in 2025, which dragged its adjusted gross margin from 40 percent down to 33 percent (The Information). Greg Brockman testified the company expects to spend on the order of fifty billion dollars on compute in 2026. Every one of those dollars is the recurring bill, not the one-time one. This is the same physical fact I traced from the hardware side in tokens per second is a memory bandwidth number: each generated token drags the weights through memory once. Inference cost is that read, multiplied by every token of every answer for every user.

The inversion, in one calculator

Here is where the rule breaks. If serving costs scale with usage, then a subscription customer who uses the product heavily can cost more to serve than they pay. In January 2025 Sam Altman said outright that OpenAI was losing money on its two-hundred-dollar-a-month Pro plan, "because people are using it much more than we expected" (TechCrunch). Not the free tier. The most expensive plan they sell. Move the sliders below and watch the sign of the profit flip as a customer falls in love with the product.

Fig. 1 · the unit-economics inversion

Pick a plan, then set how heavily this customer uses it. Illustrative cost-per-query, in the range analysts quote; the shape is the point, not the decimals.

30
$0.10
they pay$20
they cost$90

The heavier the use, the deeper the red. SemiAnalysis estimated that a fully-exercised two-hundred-dollar plan can burn up to fourteen thousand dollars of compute a year. That is not a pricing mistake anyone can fix by nudging the number, because of what happens the moment you try.

Why you cannot just raise the price

The obvious fix is to charge the heavy user more. Two forces make it a trap. The first is that the price of intelligence is in freefall. The clearest measurement comes from a16z's LLMflation analysis: GPT-3-level intelligence cost about sixty dollars per million tokens at launch in November 2021, and by late 2024 the cheapest model hitting the same score cost about six cents. A thousandfold collapse in three years, roughly tenfold every year. Scroll the curve.

Fig. 2 · your inference bill, by year

Same task, same level of intelligence, four years of prices. Set how many tokens you run a month and watch the bill for a fixed capability fall.

1B

That is the price of a fixed capability crashing. Underneath it hides the fact the chart on its own would never show you: the frontier price barely moved. GPT-3's sixty dollars per million tokens at launch in 2021 was still, roughly, the output price of a top-tier reasoning model in late 2024. The floor fell a thousandfold; the ceiling held. What changed is not the price of a token, it is how much intelligence a token at the ceiling now buys. So a lab lives between two lines: the commodity floor collapsing under its cheap tier, and a frontier ceiling that only holds as long as its best model stays ahead.

In a market where the same capability gets ten times cheaper every year, raising your price sends the customer to a competitor in the thirty seconds it takes to change an API base URL. Google has made frontier-class models free at the point of use; other labs cut prices in single announcements. You cannot charge more for a thing the whole industry is racing to give away.

The second force is the one that actually decides who survives, and it is not about pricing at all. It is about who you pay rent to.

The five-layer cake

At Davos in early 2026, Jensen Huang described the AI industry as a five-layer cake: energy at the bottom, then chips, then infrastructure, then models, then the applications on top (Forbes). It is a useful map because it shows you exactly where the money stops. For most of the last few years a frontier lab owned exactly one layer, the models. It sold up into the applications above, and it rented the three below: infrastructure, chips, and energy. Every answer ran on someone else's chips, in someone else's data center, on someone else's power. Click through the layers and see who collects.

Fig. 3 · who owns each layertap a layer

The lab owns the model layer, sells up into apps, and pays rent on the three below. The fattest rent is the chip layer.

Look at the chip layer. Nvidia runs a gross margin around 75 percent (75.0 percent GAAP in the quarter ending early 2026, from its own results). That means for roughly every dollar a lab spends on the silicon that runs inference, about seventy-five cents is Nvidia's gross profit. The model layer is bleeding while the layer directly beneath it is the most profitable large business in technology. That gap is the whole tension of the current moment: the companies capturing AI's value and the companies burning cash on AI are, for now, mostly different companies.

Everyone is burning, except the landlord

Zoom out and the pattern holds across the stack. The labs building the models post losses that scale with their revenue. The hyperscalers renting them the infrastructure are spending sums that would have been unthinkable two years ago. And underneath, the one reliably profitable layer is the one selling picks and shovels.

PlayerLayerRevenue / spendThe number
OpenAImodels~$13B revenue, 2025~$20.9B operating loss; inference cost up ~4x, gross margin 40% to 33%
Anthropicmodels~$47B run-rate, mid-2026~$3B cash burn in 2026, down from ~$5.6B in 2025; raised a $65B round at a reported ~$965B valuation
xAImodels$3.2B revenue, 2025~$6.4B net loss in 2025, reported burn near $1B a month; folded into SpaceX
Open sourcemodelsMistral ARR >$400MDeepSeek's famous $5.6M was the final training run only; real hardware spend is reported near $1.3B. Cheap to a point, not free.
HyperscalersinfraMicrosoft, Google, Amazon, Meta~$725B in combined 2026 capex, up ~77% year on year, most of it AI
Nvidiachips~$216B revenue, FY2026~$120B net income at ~71% full-year gross margin. The landlord, not a tenant.

Read the last two columns together. The frontier labs and the clouds are spending on a scale that only makes sense if the payoff is enormous and near. One chip company is turning two hundred billion dollars of revenue into a hundred and twenty billion of profit by selling the shovels. Set against this, OpenAI's roughly twenty-billion-dollar revenue run-rate is a few percent of what the big four hyperscalers alone will spend on capacity this year. That gap is the substance of what David Cahn of Sequoia called AI's 600-billion-dollar question in 2024: the revenue required to justify the buildout is far ahead of the revenue that exists. The bet is that inference gets cheap enough, and demand deep enough, to close it.

The only structural escape: own a layer you were renting

If you cannot raise prices and you cannot stop the usage, the one lever left is to own more of the cake. The most expensive rented layer is the chips, so that is where the labs are moving. In October 2025 OpenAI announced a collaboration with Broadcom to co-develop ten gigawatts of custom accelerators (OpenAI), and in June 2026 unveiled the first chip, an inference-only design, with rollout beginning in the second half of 2026. OpenAI's own claim is modest, "performance per watt substantially better than current state of the art." The louder number came from Broadcom's chief executive, who told reporters early testing showed roughly fifty percent lower cost per inference token and performance on par with Nvidia's best. That is his figure, in early lab conditions, not an audited or independent one, and it is worth holding at arm's length. But the direction is unambiguous: if inference is the bill that scales with success, owning the machine that runs it is the only move that changes the slope instead of the intercept.

Jevons has the last word

Suppose it all works. The chip lands, cost per token halves, the margin finally turns. Does the bill shrink? History says no. When a resource gets cheaper to use, people do not use less of it, they use dramatically more. Satya Nadella pointed at exactly this the week DeepSeek rattled the market in January 2025: "Jevons paradox strikes again. As AI gets more efficient and accessible, we will see its use skyrocket, turning it into a commodity we just can't get enough of" (Nadella on X).

This is why the story never ends with a stable margin. Cheaper inference does not close the bill, it enlarges the market that generates it. An agent that runs a hundred model calls to finish one task costs a hundred times a single question, and the moment agents get cheap enough, everyone runs them constantly. The bill grows because the thing it pays for got good. Every actor in this industry is straining to push the price of a unit of intelligence toward zero, and the closer they get, the more of it the world consumes. That is not a contradiction to resolve. It is the shape of the whole race, and it is worth understanding before you place any bet on who wins it.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I write about the engineering fundamentals under the AI stack: harnesses, memory, reliability, and the economics nobody itemizes.

Related: the hardware side of inference cost · More posts · X