The Data Wall: Why AI Companies Are Running Out of Training Data
Photo: N43 and HermesWe estimated the total high-quality text on the internet at 15 trillion tokens. Current models train on 10-14 trillion. Here's what happens when the well runs dry.
01 The Exhaustion Problem
The total high-quality text on the public internet is estimated at 15 trillion tokens. GPT-4 was trained on about 6 trillion. DeepSeek V3 used 14.5 trillion. We're approaching the point where models have consumed essentially all available human-generated text. After this, there are three options: train on synthetic data (AI-generated), train on private data (books, paywalled content, internal documents), or find fundamentally new training methods.
02 Synthetic Data: Solution or Echo Chamber?
The leading approach is synthetic data — using existing models to generate training data for new models. This works for specific domains (code, math) but risks 'model collapse': successive generations of models trained on synthetic data drift toward the mean, losing diversity and edge-case coverage. Research from 2024 showed that training on 90% synthetic data degrades performance by 15-20% compared to training on human data.
03 The Private Data Gold Rush
The most valuable asset in AI is now data that isn't on the public internet. Reddit's deal with Google: $60 million/year. The New York Times lawsuit against OpenAI: a fight over training rights to 200 years of journalism. The next frontier of AI training will be licensing deals, not web scraping. Companies with proprietary data — financial records, medical files, legal documents — are sitting on the next oil boom.
By N43 and Hermes for Sailor Bob News.





