Summary
AI companies face a looming "data wall," with usable human-written text expected to be exhausted between 2026 and 2032, impacting high-quality language data first. This challenge highlights a fundamental limitation: AI excels at pattern recognition but cannot create knowledge where underlying data is sparse. Experts like Adityo Prakash note that scaling models won't solve problems lacking relevant examples, as seen in drug discovery. Many organizations also lack AI-ready data, leading to project failures. Quality data consistently outperforms sheer volume. When data is insufficient, businesses should integrate AI with first-principles models or encode explicit rules from domain knowledge. The competitive edge lies in understanding data boundaries and applying complementary technologies, rather than blindly pursuing larger models. Businesses must assess their data readiness for AI tools and either generate new data or formalize existing expertise.
The largest AI companies in the world will soon face a problem they won’t be able to buy their way out of. According to research from Epoch AI, the supply of usable human-written text will be exhausted somewhere between 2026 and 2032, with high-quality language data running short first.
The industry is calling it the data wall, and it’s looming. Every proposed workaround, including synthetic generation, multimodal training and better pruning, is an attempt to manufacture a fix that used to be free.
For small business owners, the upcoming deadline isn’t so scary. They’ve been living on the other side of that wall this whole time.
The limitations of existing data
Adityo Prakash is chairman and chief executive of Verseon, a company that designs drug candidates by pairing AI with physics-based molecular modeling. His industry hit the data ceiling before anyone was calling it a wall, because there is a huge gap between what has been tested and what exists.
"One of the biggest misconceptions about AI is that increasing the size of a model automatically allows it to solve problems where the underlying data doesn't exist," Prakash said in a statement.
"AI is extraordinarily good at learning patterns from information it has seen, but there are scientific and business problems where the relevant examples are simply too sparse. No amount of scaling can magically create experimental knowledge that was never generated."
Drug discovery is an example of an industry where AI cannot simply learn its way past gaps in the underlying data.
"The universe of possible drug-like molecules is unimaginably larger than the fraction humanity has ever synthesized and tested," Prakash said. "If your AI system learns primarily from known molecules, you have to ask whether you're actually discovering something new or becoming increasingly sophisticated at navigating what we already know."
That’s not an insignificant problem. A model trained on what a business already knows will get very good at restating it, and that’s probably not the answer that’s going to move an enterprise forward at the pace needed to keep up.
‘The data is not ready for AI’
Gartner found that 63% of organizations either do not have or are unsure whether they have the right data management practices for AI, based on a July 2024 survey of data management leaders. The firm further predicted that organizations would abandon 60% of AI projects unsupported by AI-ready data through 2026
There’s a mismatch between how companies store data and what AI requires of it, said Roxane Edjlali, a senior director analyst at Gartner
"Traditional data management operations are too slow, too structured, and too rigid for AI teams," she said, noting that in conventional practice, uses of data are poorly documented and data sits siloed across systems.
Her summary of the readiness test is one sentence long. "If the data has issues, then the data is not ready for AI."
Consider how AI is sold to a small business. The pitch is that the tool learns your business, but most organizations with dedicated data teams aren’t equipped to meet even that level of commitment.
For small businesses, the process is difficult. Forty customers is not a training set. Eighteen months of invoices is not a pattern. A founder who has priced every job by instinct has decades of judgment and no dataset, and a model pointed at that business will generalize confidently from what turns out to be almost nothing.
Prakash frames the diagnostic as a question that often gets skipped.
"Do we actually have the data required for this problem?" he said. "If the answer is no, selecting a larger model may not solve anything. You may need domain knowledge, simulations, first-principles models, new data generation, or an entirely different technical architecture."
Bigger has not reliably beaten better
This information is not new. Percy Liang, a computer science professor at Stanford, put it plainly to MIT Technology Review in 2022, when scaling was still the consensus answer to everything.
"We've seen how smaller models that are trained on higher-quality data can outperform larger models trained on lower-quality data," he said.
The Epoch AI research highlights how the line between high-quality and low-quality data is itself fuzzy, and that researchers filter aggressively toward the higher-quality information precisely because it is the language they want reproduced.
Quality has always done more work than volume, and the data wall is what happens when the that high-quality pile dwindles down.
What you use instead
Verseon's answer is physics. When you cannot learn a molecule's behavior from thousands of prior examples, you model the forces that govern how molecules interact.
"This isn't an argument against AI. It's an argument for understanding what AI is good at and where complementary approaches are necessary," Prakash said. "We combine AI with physics-based modeling because physics gives you a way to reason about molecular interactions even when you don't have thousands of historical examples available for a particular type of molecule."
No small business is running molecular simulations, but that doesn’t mean there aren’t lessons to be learned. The transferable move is structural: When pattern-matching has nothing to match against, supply the rules instead.
That means encoding what you know to be truehe pricing floor below which a job is not worth taking. The three conditions that have always predicted a client will churn. The steps of a process that cannot be reordered.
Written down, those are the constraints a system can operate inside. Left only in a founder's head, they are the reason AI does not know your margins and is therefore not making you money.
The counterargument
The data wall may not arrive as forecast. Recent advancements have extended runway, and Epoch's own analysis treats continued scaling through 2030 as plausible but not an inevitability.
There is also a self-interest problem. A company selling physics-based modeling has an obvious reason to argue that data-driven AI has limits, and Prakash's argument should be weighed with that in mind. Gartner sells data management research. Both stand to gain from a doomsday scenario.
What to do this quarter
Take three decisions where someone has suggested using an AI tool. For each, write down what information the AI would need to make a good decision, , then check whether you have that data in a form AI could read.
Where the answer is yes, proceed. Where it is no, the fix is not a different or even a better tool. You either need to start generating the data deliberately, which takes months and should start now, or turn the judgment you already have into explicit rules, which takes an afternoon.
The distinction matters because most AI investments that fail were never going to work regardless of which model got chosen. They were aimed at questions the business had no evidence to answer.
"The real competitive advantage will come from understanding where your data ends, what your models can legitimately infer from it, and what additional knowledge or technology is required to cross that boundary," Prakash said.
The frontier labs are about to spend billions working out what to do when the data runs out. A small business owner can answer the same question this week, on one page, for free.
