Skip to content
Blog

AIData

Model Collapse: What Happens When AI Learns From AI

Models trained on AI-generated data degrade — losing rare knowledge, amplifying errors, drifting from reality. We break down the mechanics and consequences.

Dmitry3 min read

A Photocopier Copying a Photocopy

Imagine you photocopy a document. Then copy the copy. Then copy the copy of the copy. After twenty iterations, the text is still legible, but images have turned to grey blobs, fine details have vanished, and artifacts have compounded.

This is exactly what happens to language models trained on data generated by other language models. Researchers call it model collapse.1

The Mechanics of Degradation

In 2023, a group of researchers from Oxford published a paper with a simple but troubling conclusion: sequential training on AI-generated content systematically destroys models.

Degradation occurs in two stages:

Early collapse — the model begins ignoring rare, atypical data. Unlikely but real events are "washed out" of the distribution. The model becomes more "average."

Late collapse — diversity of outputs drops sharply. In extreme cases, it produces the same response to different questions.

Loading diagram…
Knowledge distribution degradation across AI training generations

The Scale of the Problem

The internet is the primary source of training data for large language models. By various estimates, by 2026 more than half of all public text content will have been created or significantly reworked with AI assistance.

ProblemConsequence
Erosion of rare knowledgeReduced reliability in non-standard situations
Error amplificationArtifacts from previous models become entrenched
Loss of diversityHomogenization of language and style across the internet
Accumulated biasSystematic errors compound with each generation

Why Rare Data Matters Most

The paradox is that "rare" data is often the most valuable.2

Diagnosis of critical medical conditions is rare — but the quality of that knowledge determines whether a patient lives. Unusual legal cases are rare — but they demand the greatest precision. Information security incidents are rare — but the cost of a mistake is catastrophic.

A model that knows the "average" well and the "edges" poorly has an unacceptable risk profile for critical applications.

What This Means for Enterprise AI

Key insight

A company's proprietary data is not contaminated by AI generation. This makes it more valuable every year — as open internet sources continue to degrade.

Proprietary data is protection against collapse. Internal data — documentation, operational records, historical cases — is a "clean" distribution of real events, including rare and non-standard ones.

Synthetic data requires caution. Synthetics are useful for augmentation — but shouldn't form the foundation of a corpus.

Data versioning matters more than code versioning. If you don't know what percentage of your training data was created or reworked by AI systems — you don't control model quality.

Potential Solutions

Strategies for Protecting Against Model Collapse

Loading diagram…

Conclusion

Model collapse isn't a hypothetical threat. It's a mathematically predictable consequence of training on AI-generated data. The internet as a source of training corpora is degrading faster than we can fully grasp.

For companies, this means one thing: proprietary data isn't just an asset — it's a survival condition in the AI era. Organizations that accumulate, structure, and protect their data today will, in a few years, own something that can't be purchased from OpenAI or anyone else.

Footnotes

  1. Shumailov I. et al. "The Curse of Recursion: Training on Generated Data Makes Models Forget" (2023). Oxford University and co-authors demonstrated that after just a few generations of recursive training, models lose the tails of their distribution and degrade toward the "average." ↩

  2. This is known in statistics as the "tyranny of the majority": systems optimized for the median systematically fail at the tails of the distribution — precisely where the cost of error is highest. ↩

Tags:AIData

Author

Dmitry

Found it useful? Share it with your colleagues