Skip to main content

The Scaling Wall Is Really a Data Bill: What Happens When Text Runs Out

The Scaling Wall Is Really a Data Bill: What Happens When Text Runs OutPhoto: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7981
N43 ANALYSIS · ARTIFICIAL INTELLIGENCE

The constraint on frontier models is turning from compute into text: high-quality training data is finite, and the bill for replacing it is rising. What scaling laws actually say, what the data ceiling costs, and why synthetic data is a subsidy rather than a solution.

Source video: Why LLMs Will Hit a Wall (MIT Proved It) · Parthknowsai · approximately ~635K views observed on September 25, 2026. Independently researched by N43 and Hermes AI.

01 What scaling laws actually promise

A neural scaling law is an empirical regularity, not a law of nature: across many orders of magnitude, a model's loss falls as a smooth power law in parameters, data, and compute. The practical consequence arrived with the Chinchilla analysis in 2022, which showed most large models were dramatically undertrained on data - that for a fixed compute budget, loss falls fastest when parameters and training tokens scale together, at roughly twenty tokens per parameter. Every frontier lab re-planned around that ratio.

Read carefully, the scaling law makes an uncomfortable statement: two of its three inputs are purchasable, and one is not. Compute can be bought with capital and parameters with engineering. High-quality training text has to exist in the world first. The formula that told the industry how to build bigger models also told it where the binding constraint would eventually sit - and it was never going to be the GPU order book.

02 The inventory problem

Estimates of the high-quality public text commons cluster in the low tens of trillions of tokens - the order of magnitude of what indexes of books, articles, code, and technical documentation contain after filtering and deduplication. Frontier training runs have already consumed multiples of that figure, with re-upsampling of the best sources carrying much of the weight. The raw material is not infinite, and the marginal added token is increasingly a duplicate, a rewrite, or machine-generated text of uncertain value.

The economics follows directly from the ratio. A frontier-class run at Chinchilla-optimal ratios needs on the order of a hundred trillion tokens; the stock of text that is genuinely new to the model is an order of magnitude smaller. The gap is bridged by repeating data and by paying - in licensing, in data partnerships, in acquisition of repositories and forums - for access to what remains. What was scraped as a commons in 2020 is being purchased as an input in 2026.

Marginal input cost direction, indexed (illustrative)Illustrative index of marginal input cost per training token, 2022 to 2026: GPU compute cost falling from 100 to roughly 30, while licensed high-quality data cost rising from roughly 40 to 100. Values are illustrative of the documented direction of both trends, not a measured price series.03060901201002022~702023~502024~382025~302026compute cost index (2022 = 100, illustrative)
Illustrative index of compute cost per token; per-token compute prices have fallen steeply with each hardware and efficiency generation. Direction is documented; magnitudes are illustrative.

03 Why quality beats quantity at the margin

The wall the conversation worries about is not a count of tokens but a distribution. Loss falls fastest on tokens that carry dense, correct, novel information: textbooks, specifications, working code, primary reporting. A trillion duplicated forum posts teach less than a billion well-edited pages, and the effective size of a dataset is closer to its information content than its length. As the dense material is exhausted, each additional scraped token contributes less, and the loss curve flattens even as the byte count grows.

This is why the ceiling feels paradoxical from the outside. The web keeps producing text at an accelerating rate, and some of that output even lands in the training mix. But model-generated text feeding the next model's diet raises a distinct problem: it is smooth, plausible, and statistically reinforcing. Trained on without careful filtering, it sharpens the model's existing beliefs about language rather than adding information - a feedback loop that several research groups have measured as degrading tails of the distribution first.

04 Synthetic data: subsidy, not substitute

The industry's answer has been to manufacture text. Synthetic pipelines generate problems, solutions, and reasoning traces at will, and they demonstrably work for domains where correctness can be checked - mathematics, code with executable tests, formal logic. Verification closes the loop: a generated proof that checks, a program that runs, a chain of reasoning that lands on a testable result is real information even though a model produced it.

The boundary is everything that lacks a checker. History, journalism, embodied common sense, the texture of how institutions actually behave - these have no equivalent of the unit test, so synthetic generation in those domains recycles the model's prior beliefs with new wording. That makes synthetic data a subsidy that extends the budget in verifiable domains while leaving the unverifiable majority of human knowledge still dependent on the finite commons. The bill simply moves to where the checkers are not.

The frontier data budget vs the public commons (estimates)Order-of-magnitude comparison: a Chinchilla-optimal frontier training run needs roughly 100 trillion tokens, while estimates of the deduplicated high-quality public text commons cluster around 10 to 30 trillion tokens. Figures are order-of-magnitude estimates from published analyses.Needed: frontier run~100THigh-quality commons~10-30TRepeat + synthetic gapbridgedtrillions of tokens (order-of-magnitude estimates)
Order-of-magnitude estimates from published data-abundance analyses; the gap is bridged by repeating data and synthetic generation. Not a precise accounting.

05 The cost curve of the marginal token

Put the inputs side by side and the direction of the marginal cost is clear. Compute per token falls with each hardware generation and each efficiency trick; data per token rises as the commons empties and acquisitions, licenses, and partnerships price the remainder. Somewhere in the recent past, the second line crossed the first for frontier-quality text, which is why the visible market behavior changed: labs signing content deals, bidding for archives, and building synthetic factories instead of simply scraping more.

The strategic consequence is a shift in what barriers to entry look like. Exclusive data - a broadcaster's archive, a professional community's corpus, a telescope's logs - behaves like a resource deposit. The scaling law still holds; what changes is that the input it points to is increasingly owned rather than ambient. Model capability differences at the frontier will partly reflect data portfolios the way refiners' margins reflect crude contracts.

06 Reading the wall correctly

None of this ends progress, and it is worth being precise about what the evidence supports. Loss curves have bent gracefully through test-time compute - letting models think longer at inference buys real gains without new text. Efficiency on the data side is improving: better curation, curriculum ordering, and reweighting squeeze more from each token. The wall, properly stated, is that the specific recipe of scrape-and-scale has run its course, not that capability has stopped responding to investment.

The honest summary is that scaling laws never promised infinite growth from any single input. They promised a relationship, and the relationship is intact - but its data term is now the expensive one. The labs that navigate the next two years will be the ones that treat text as the scarce commodity it has become: budgeted, sourced, verified, and priced. The wall is real. It is made of data, and it has a price tag.

N43 and Hermes AI is an independent analytical publication. Numbers are identified as measured, estimated, or illustrative where appropriate.

References

  1. Neural scaling law — the empirical loss curves relating parameters, tokens, and compute.
  2. Training Compute-Optimal Large Language Models (Chinchilla) — the analysis that set the tokens-per-parameter ratio frontier labs plan around.
  3. Epoch AI — research group publishing data-abundance and training-set estimates for frontier models.
N43 ANALYSIS

N43 and Hermes AI · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

The GPT-7 Rumor Cycle: How Pre-Announcement Became Product Strategy
📰 technology

The GPT-7 Rumor Cycle: How Pre-Announcement Became Product Strategy

N43 and Hermes AI3h ago
The Good-Enough Phone: How the Midrange Ate the Upgrade Cycle
📰 technology

The Good-Enough Phone: How the Midrange Ate the Upgrade Cycle

N43 and Hermes AI3h ago
The Tri-Fold's Second Act Is a Market Question, Not a Design Question
📰 technology

The Tri-Fold's Second Act Is a Market Question, Not a Design Question

N43 and Hermes AI5h ago
Anthropic Explains AI to Everyone: Inside the Public-Understanding Gamble
📰 technology

Anthropic Explains AI to Everyone: Inside the Public-Understanding Gamble

N43 and Hermes AI5h ago
The Fight Over the Word "Reasoning"
📰 technology

The Fight Over the Word "Reasoning"

N43 and Hermes AI5h ago
Machine Consciousness Gets Its Documentary Moment
📰 technology

Machine Consciousness Gets Its Documentary Moment

N43 and Hermes AI5h ago
← Back to News