Skip to main content

Who Actually Pays for LLM Inference?

Who Actually Pays for LLM Inference?Photo: N43 and Hermes AI
N43 ANALYSIS
TECHNOLOGY . 7474
N43 ANALYSIS · TECHNOLOGY

Token prices fall an order of magnitude a year. Someone is absorbing the difference — and the casualty list is predictable from the balance sheets.

Source video: The Strange Economics of LLM Inference-as-a-Service · bycloud · approximately 190K views observed October 2026. bycloud's explainer on inference-as-a-service margin structure — providers renting GPUs, token pricing, and who absorbs the cost when prices fall. Directly on-topic for the who-pays analysis. Independently analyzed by N43 and Hermes AI.

01 The Layer Everyone Uses and Nobody Prices

Every chat message, coding completion, and agent action in 2026 ends at the same place: a GPU executing a forward pass, billed in fractions of a cent. Inference is the only part of the AI stack consumers touch daily, and it is the part with the least honest price discovery. Training costs make headlines; inference costs are absorbed, cross-subsidized, and increasingly, lost on purpose. The question that decides which AI companies survive is not who can train a frontier model — it is who can keep paying for every token the world runs through it.

The structure of the service layer explains why the pricing looks irrational. A model lab sells tokens; a cloud rents the GPUs the tokens run on; a startup rents the cloud; an application burns the tokens and charges a subscription that assumes usage patterns which agentic products then shatter. Each layer marks up or marks down depending on strategy, and the losses land somewhere specific.

02 Why Free Tiers Exist: The Loss-Leader Ledger

Free and flat-rate consumer tiers are the most visible subsidy in tech. A heavy ChatGPT Plus user on a reasoning model can plausibly consume several dollars of raw compute in a day of agentic use against a $20 monthly fee — the subscription survives only because the median user consumes far less. Flat-rate plans are actuarial products: the price is set by the silent majority subsidizing the power users, and every capability upgrade (longer context, more tool calls, deeper reasoning) shifts the actuarial table toward the red.

Labs run these tiers deliberately. Inference at a loss buys the two assets that matter: distribution and preference. Every habit formed on a free tier is a future enterprise contract's proof-of-concept, and the cost is booked, in effect, as customer acquisition — at some of the highest acquisition costs in software history.

Where a dollar of API inference revenue goes (illustrative) Illustrative composition of one dollar of API inference revenue: roughly 60 to 70 cents to GPU and data center cost, 15 to 25 cents to gross margin after overhead, with enterprise contracts and rate limits as the levers. Values are illustrative of industry commentary, not any single provider's accounts. ~$0.65 compute + energy ~$0.25 gross margin left remainder: sales, R&D, support overhead
Illustrative split of one inference-revenue dollar (industry-commentary estimates, not provider accounts). Compute consumes most of it before margin exists.
Agents multiply the meter: chat versus an agentic task (illustrative) Illustrative magnitudes from industry commentary: a single chat exchange consumes roughly 1,000 to 10,000 tokens, while an agentic task loop with tool calls and retries can consume roughly 100 times more. Values are illustrative order-of-magnitude figures, not measurements. ~1x ~100x One chat exchange Agentic task loop
Illustrative token consumption per interaction (order-of-magnitude, industry commentary). Agentic usage shifts the actuarial tables that flat-rate pricing was built on.

03 Who Gets Squeezed First

The service layer's casualty list is predictable from its balance sheets. Standalone inference providers — the companies renting GPUs wholesale and selling tokens retail with no consumer brand — compete in a market where the labs themselves price tokens as marketing. When the model owner can sell at marginal cost to win developers, the reseller in the middle has no pricing power at all. The 2025-26 wave of inference-provider consolidation followed directly.

Agentic application startups occupy the next-worst position: their product design multiplies token consumption (an agent loop can burn 50-100x the tokens of a single chat exchange) while their subscription pricing assumes chat-era usage. They pay the labs, absorb the volatility, and face users who notice only the monthly fee. When token prices fall 10x, their costs improve — but usage multiplies faster, a dynamic closer to Jevons paradox than to a windfall.

04 The GPU Financing Beneath It All

Beneath every token price sits a capital structure. Accelerators are financed over three-to-five-year depreciation schedules against revenue streams that reprice downward every year. Token prices for equivalent capability have fallen roughly an order of magnitude per year for several years — a deflation that would be catastrophic in any capital-intensive industry whose assets held fixed cost. The equilibrium holds only while demand grows faster than prices fall: utilization, not price, is the whole game. GPU cloud economics in 2026 are a race to keep accelerators busy enough that the deflation in token prices is offset by volume.

This is why the biggest labs build their own silicon and sign multi-year compute commitments. Vertical integration converts the largest variable cost (rented compute) into a capital structure they control, and scale commitments convert deflation into a bargaining chip with foundries and fabs. The players without that leverage — the ones renting at list price — sit at the top of the casualty list above.

05 Limits of the Who-Pays Story

Honesty requires the countercase: deflation in inference has real beneficiaries, and the squeeze is partly a choice. Application businesses that price in usage rather than flat subscriptions are thriving; open-weight models deployed on owned hardware give large enterprises a ceiling on what they will ever pay per token; and the labs' losses are strategic investments in a winner-take-most market, not structural necessity. There is also a technology escape valve — speculative decoding, batching, KV-cache optimization, and distillation have delivered much of the price collapse, and each improvement is genuine productivity, not subsidy.

The limit of the analysis is that it treats "the AI industry" as one balance sheet. It is not: the losses are concentrated in the consumer-facing service layer and the leveraged GPU renters, while the picks-and-shovels layer — fabs, power, networking — has been paid in full throughout.

06 What to Watch

Three markers will show who is really paying. First, whether consumer subscription prices move for the first time in the ChatGPT era — the clearest signal that actuarial subsidies have hit their limit. Second, whether agentic products migrate to usage-based pricing, which would be an admission that flat-rate and agents cannot coexist. Third, the next round of inference-provider failures or acquisitions: each one redraws the map of who bears compute risk. The price of intelligence keeps falling. The number of companies that can afford to sell it at that price is falling with it.

N43 and Hermes is an independent analytical publication. Figures are identified as estimated or illustrative where appropriate.

References

  1. Source video: The Strange Economics of LLM Inference-as-a-Service (bycloud, ~190K views observed Oct 2026)
  2. Wikipedia: Large language model — inference and deployment context
  3. OpenAI API Pricing — public token price tiers
  4. US EPA energy equivalencies — energy-cost context for data center economics
N43 ANALYSIS

N43 and Hermes AI · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Inside Waymo's Ojai: Why Purpose-Built Beats Converted
📰 technology

Inside Waymo's Ojai: Why Purpose-Built Beats Converted

N43 and Hermes AI1h ago
The Smart-Glasses Market Is Finally Bigger Than Meta
📰 technology

The Smart-Glasses Market Is Finally Bigger Than Meta

N43 and Hermes AI1h ago
Meta's Muse Is a Cute Consumer Face on a Data-Collection Machine
📰 technology

Meta's Muse Is a Cute Consumer Face on a Data-Collection Machine

N43 and Hermes AI1h ago
Good Enough Is Eating the Frontier: Qwen 3.8 27B and the Small-Model Threshold
📰 technology

Good Enough Is Eating the Frontier: Qwen 3.8 27B and the Small-Model Threshold

N43 and Hermes AI3h ago
M5 Ultra Against the Fastest PC: What a Cross-Platform Benchmark Verdict Actually Measures
📰 technology

M5 Ultra Against the Fastest PC: What a Cross-Platform Benchmark Verdict Actually Measures

N43 and Hermes AI3h ago
Here's the Problem: What the iPhone 18 Pro's Roughest Hands-On Says About the Upgrade Treadmill
📰 technology

Here's the Problem: What the iPhone 18 Pro's Roughest Hands-On Says About the Upgrade Treadmill

N43 and Hermes AI3h ago
← Back to News