AI Famine Crisis: The granary is full, yet the rice remains unwashed.
At the start of 2025, Elon Musk said something that sent a chill down the entire industry's spine during a livestream: We have essentially exhausted the cumulative sum of human knowledge.
In the same year, Ilya Sutskever — the person who personally elevated the Scaling Law to its iconic status — turned around and stated: The 2010s were the era of scaling. Now we have returned to the era of wonder and discovery. Everyone is searching for the next breakthrough point.
By 2026, this debate has gained a unified name: the Data Wall.
He Baohong, Chief Engineer of the China Academy of Information and Communications Technology, put it more bluntly at a conference in May: Model pre-training on the internet has hit the Data Wall. Public domain data has been fully depleted, and there are barely any new data sources left to effectively boost model performance.
It sounds alarming. AI labs all over the world are competing for the same shrinking pool of raw materials, just like a group of miners who have reached the end of a mineral vein.
But after working on the front lines of data for so many years, my first reaction to these reports was not panic, but a highly inappropriate question:
In 2025, the total volume of data generated in China in one year reached 52.26 ZB, a year-on-year increase of 27.28%, accounting for 27.44% of the total global data volume.
What does 52 zettabytes mean? 1 ZB is equal to 10 to the power of 21 bytes. This figure is still expanding at a rate of over a quarter every year.
On one hand, data is exploding at a rate of dozens of ZB per year; on the other, model vendors are complaining that they have nothing to feed their models.
This is not called famine. This is called the granary being full, but no rice has been washed for cooking.
What AI is running out of is not data, but data that can be directly put into the pot for use.
The gap between these two concepts is the entire reason for the existence of the craft of data governance.
I. The Grain Is Gone: Three Figures
Don't rush to refute. Let's lay out all the evidence supporting the "grain shortage theory". These claims are not alarmist, they are based on calculations.
The first figure: 300 trillion
Research from Epoch AI provides a widely cited estimate: the total stock of high-quality, publicly accessible human text on the internet is approximately 300 trillion tokens.
This pool contains Wikipedia, open source code repositories, academic journals, digital libraries, and forum discussions accumulated over more than 20 years. It sounds vast — but it is finite.
Over the past three training cycles, cutting-edge models have almost completely consumed this resource. Every digitized book, every public GitHub repository, every Reddit post from 2008 to 2022, and every open access medical paper has been converted into neural network weights.
Yet the speed at which humans write new books and new papers is far from keeping up with the exponentially growing appetite of next-generation training clusters.
The second figure: 25%
If the limited stock of data is a natural disaster, then this next point is a man-made disaster.
The MIT-led "Data Provenance Initiative" conducted a study examining changes in three mainstream training datasets: C4, RefinedWeb, and Dolma. The conclusion is: In just one year, 5% of all data and 25% of the highest-quality sources have been placed under access restrictions.
The New York Times, Reddit, Stack Overflow, and a large number of news agencies have either modified their terms of service or directly deployed technical blockades.
Note that 25% — what has been locked away is not marginal content, but precisely the highest quality portion. It is like the granary has not run out, but the best few bags of grain have been locked up by others.
The third figure: 2026–2032
Epoch AI's calculations suggest that high-quality human-generated text will be completely exhausted between 2026 and 2032, with the exact timeline depending on how aggressively labs pursue "overtraining".
Interestingly, this time window has been shifting backwards. Their earlier estimates were more pessimistic, but the timeline has since been pushed back. This revision itself illustrates one thing: The position of this wall is calculated, not stumbled upon by accident.
An earlier paper published in 2022 once predicted that high-quality English text would be exhausted in 2026. Looking back from 2026, the accuracy of this prediction is indeed unsettling.
II. It Is Not That There Is No Grain, But That The Grain Cannot Be Accessed
Now, please allow me to expand on that inappropriate question I raised earlier.
The gap between 52.26 ZB and 300 trillion tokens is not just a difference in order of magnitude, but a difference in nature.
Let's do a rough conversion to get a sense of scale. One Chinese character roughly corresponds to 1.5 to 2 tokens. Taking the median value, 300 trillion tokens are roughly equivalent to 180 trillion Chinese characters. That sounds like a lot, but 1 ZB is 10 to the power of 21 bytes — even if we only count text, the data volume of dozens of ZB far exceeds this figure.
So the question arises: Where did the rest of the data go?
The answer is — almost all of it is locked in three places.
The first place: Enterprise private domains. How many years of order records have been running in your ERP, how many real calls have been stored in your customer service system, how many fault analysis reports have been accumulated by your R&D department. This data has never been on the internet, and cannot be captured by web crawlers. They are the largest batch of "unexploited farmland" on this planet.
The second place: Government and institutional internal systems. Healthcare, justice, education, meteorology, natural resources. These data sets are huge in volume and extremely high in quality, but for compliance and sovereignty reasons, the vast majority will not be opened to commercial model training.
The third place: The physical world. An article in the *Learning Times* includes a very illustrative comparison: The scale of training corpora that supports general large models to achieve capability leaps reaches the order of trillions of tokens, while the high-quality real machine interaction datasets available globally for embodied intelligence training are generally still at the scale of millions of trajectories. The difference spans multiple orders of magnitude.
So the truth is: There has never been more data on Earth than there is today, yet the small fraction of data that models can access legally, at low cost, and with high quality is running out.
This is not resource depletion. This is structural failure on the supply side — the grain is in the granary, but there is no channel to turn it into rice that can be cooked.
III. The Internet Is Being Polluted By Its Own Excreta
There is one more thing more troublesome than "limited stock" that many people have not realized.
Between 2023 and 2025, millions of auto-publishing bots, SEO content farms, and affiliate marketing tools poured billions of pages of low-cost machine-generated text into the public internet.
By 2026, if a lab crawls the open web, the content it captures will most likely be synthetic text generated by last year's batch of models.
There is a harsh but accurate saying: The AI industry has polluted its own water source.
What are the consequences? Researchers now spend more engineering time building filters to identify and discard machine-generated text than they spend working on the actual model architecture.
This creates a particularly absurd situation —
To bypass the Data Wall, the industry uses AI-generated content on a large scale; and the proliferation of AI-generated content submerges the remaining real data even deeper. The more urgently you try to solve the problem, the more expensive the problem becomes.
This has led to the most bizarre enclosure movement in commercial history: Verified, clean text written by real people before 2022 has become digital gold.
Tech companies are spending hundreds of millions of dollars to acquire closed private human archives — decades of medical records from hospital networks, customer service call records from telecom giants, and physical archives scanned by historical societies. Publishers, newspaper groups, and forum platforms, once seen as "leftover relics of traditional media", have suddenly gained huge bargaining power in licensing negotiations.
If you hold 20 years of forum discussions written by real people, talking about real faults in the real physical world — the thing you are sitting on is an irreplaceable training asset that cannot be artificially manufactured.
Is Synthetic Data The Savior? Half Yes, Half No
When there is not enough grain, people will try to make their own, which is an instinctive reaction. How large is the scale of the industry's response? Look at two figures:
According to estimates, 30% to 60% of the training tokens in recent cutting-edge model training are synthetic, not crawled from the web.
Gartner predicts that by the end of 2026, synthetic data will account for about 75% of all data used in AI development — compared with roughly 1% in 2021.
That is a 75-fold increase in five years. This curve is steeper than the iteration pace of any model.
But the cost is: Model Collapse
The landmark study published in *Nature* clearly demonstrated what happens with recursive training:
For the first generation, minor statistical errors and edge cases are smoothed out; by the third generation, the model begins to forget rare facts, uncommon linguistic nuances, and the perspectives of minority groups; by the fifth generation, the output degenerates into repetitive, low-variance gibberish.
The mechanism is actually easy to understand. A language model is a probability prediction engine that samples from the center of the human distribution, preferring common words and typical structures. When you train a model with synthetic data, you are training with an "approximation of an approximation" — the tail of the probability distribution disappears.
This process is similar to inbreeding in biology: each generation amplifies subtle defects a little more.
That is why mainstream teams have reached a highly consistent ratio: Synthetic data accounts for 30% to 50%, and the rest must be real and manually curated. Never train with pure synthetic data, and always keep a small set of human-generated seeds as anchors.
The opposing side also has a valid point: AlphaGo won exactly this way
But it is also wrong to dismiss synthetic data entirely. There is one counterexample too strong to ignore.
AlphaGo first learned from human chess records, then played millions of self-play games, and eventually made moves no human had ever played before. It surpassed humans precisely because it was no longer limited by human data.
Extend this logic to mathematics and programming — AI can generate new math problems and solutions, verify whether the solutions are correct, and then train with these verified synthetic data. Since mathematics and code have objective standard answers, the volume of effective training data becomes nearly infinite.
SynthLLM from Microsoft Research Asia works exactly this way: it uses graph algorithms to recombine high-level concepts from existing corpora into new synthetic samples, specifically to reduce path dependence on web crawling.
So who is right in the end?
The bottleneck of synthetic data has never been the generator, but the verifier.
In any field with objective evaluation criteria — whether code can be compiled, whether a mathematical formula holds, whether a chess game is won or lost — self-play works and can expand infinitely. In any field that relies on human judgment — creative writing, subtle nuance handling, aesthetics and tradeoffs — synthetic data will only accelerate degradation, because no one can reliably tell the model "this one is better".
In other words: You can only synthesize as much data as you can verify. And an enterprise's private domain data is essentially a ready-made verifier.
China's Solution: Do Not Look For New Grain, Wash The Old Grain Clean
This section is the part I most want to write. Because when facing the same Data Wall, the response logic on both sides is completely different.
The mainstream narrative in Silicon Valley is "looking for new grain" — buying copyrights, making synthetic data, betting on multimodality, betting on embodied intelligence.
The answer given by China is another approach: Wash the existing grain clean.
On June 3, 2026, the National Data Bureau issued the *Implementation Plan for Promoting High-Quality Industry Dataset Construction Action* (National Data Science and Technology Foundation Document No. 25 [2026]). This document contains a very precise definition — a high-quality industry dataset is "a collection of industry data that has undergone data processing such as collection and processing, can be directly used to develop and train artificial intelligence models, and can effectively improve model performance".
Please pay attention to the four words "can be directly used". It confirms what I said in the second section: There is a processing procedure between raw data and usable data.
Six Special Actions
The plan deploys six initiatives, which I have rearranged logically according to my own understanding:
The last line deserves extra explanation. The plan clearly proposes to "explore new transaction modes such as token trading, and build a data value system based on tokens that can be quantified and priced".
The significance of this statement is that The unit of measurement for data is changing from "number of entries" and "GB" to tokens actually consumed by models. Selling data by the number of entries and pricing by tokens actually called by agents are two completely different business models. This is two sides of the same coin as the first token transaction I saw at the Yunqi Conference.
In one quarter, the volume grew from 960 PB to 1565 PB.
In the same period, the National Dataset Management Service System was officially released and launched for trial operation, with the goal of building a national-level management system that is "physically decentralized but logically centralized".
Looking at the three solution paths proposed by He Baohong from the China Academy of Information and Communications Technology, they fully align with this national plan: The first is to move from public domains to private domains, to deeply develop and utilize private domain data for specific industries and scenarios; The second is to develop synthetic data, but use it with controlled volume; The third is to improve data quality, and continuously optimize the quality of existing data with advanced data engineering methods.
Two of the three paths point to the same thing — washing rice.
Six. Three Things Enterprises Should Do
After talking so much about macro-level issues, let's come back to your own enterprise.
If the "grain shortage" judgment holds, its meaning for different roles is completely opposite. For model vendors it is a crisis; for enterprises that hold industry data, it is a revaluation of assets.
The first thing: Separate the calculation of "having data" and "having AI-ready data"
This is the step