A Photocopier Copying a Photocopy
Imagine you photocopy a document. Then copy the copy. Then copy the copy of the copy. After twenty iterations, the text is still legible, but images have turned to grey blobs, fine details have vanished, and artifacts have compounded.
This is exactly what happens to language models trained on data generated by other language models. Researchers call it model collapse.1
The Mechanics of Degradation
In 2023, a group of researchers from Oxford published a paper with a simple but troubling conclusion: sequential training on AI-generated content systematically destroys models.
Degradation occurs in two stages:
Early collapse — the model begins ignoring rare, atypical data. Unlikely but real events are "washed out" of the distribution. The model becomes more "average."
Late collapse — diversity of outputs drops sharply. In extreme cases, it produces the same response to different questions.
The Scale of the Problem
The internet is the primary source of training data for large language models. By various estimates, by 2026 more than half of all public text content will have been created or significantly reworked with AI assistance.
| Problem | Consequence |
|---|---|
| Erosion of rare knowledge | Reduced reliability in non-standard situations |
| Error amplification | Artifacts from previous models become entrenched |
| Loss of diversity | Homogenization of language and style across the internet |
| Accumulated bias | Systematic errors compound with each generation |
Why Rare Data Matters Most
The paradox is that "rare" data is often the most valuable.2
Diagnosis of critical medical conditions is rare — but the quality of that knowledge determines whether a patient lives. Unusual legal cases are rare — but they demand the greatest precision. Information security incidents are rare — but the cost of a mistake is catastrophic.
A model that knows the "average" well and the "edges" poorly has an unacceptable risk profile for critical applications.
What This Means for Enterprise AI
A company's proprietary data is not contaminated by AI generation. This makes it more valuable every year — as open internet sources continue to degrade.
Proprietary data is protection against collapse. Internal data — documentation, operational records, historical cases — is a "clean" distribution of real events, including rare and non-standard ones.
Synthetic data requires caution. Synthetics are useful for augmentation — but shouldn't form the foundation of a corpus.
Data versioning matters more than code versioning. If you don't know what percentage of your training data was created or reworked by AI systems — you don't control model quality.
Potential Solutions
Strategies for Protecting Against Model Collapse
Conclusion
Model collapse isn't a hypothetical threat. It's a mathematically predictable consequence of training on AI-generated data. The internet as a source of training corpora is degrading faster than we can fully grasp.
For companies, this means one thing: proprietary data isn't just an asset — it's a survival condition in the AI era. Organizations that accumulate, structure, and protect their data today will, in a few years, own something that can't be purchased from OpenAI or anyone else.
Footnotes
-
Shumailov I. et al. "The Curse of Recursion: Training on Generated Data Makes Models Forget" (2023). Oxford University and co-authors demonstrated that after just a few generations of recursive training, models lose the tails of their distribution and degrade toward the "average." ↩
-
This is known in statistics as the "tyranny of the majority": systems optimized for the median systematically fail at the tails of the distribution — precisely where the cost of error is highest. ↩


