Skip to main content
AI Data

Is It Safe to Train AI Models on AI-Generated Data?

July 2026 · 8 min read · Updated September 2026

Short answer. Conditionally, and the condition does most of the work. A 2024 Nature study showed that indiscriminately training models on recursively generated content causes model collapse — rare cases disappear first, then output narrows overall. Synthetic data used deliberately, mixed with real data, filtered on quality and anchored to human-verified material is a normal, effective technique. Synthetic data accumulated accidentally and unlabelled is the failure case most organisations are exposed to without choosing it.

Key takeaways

  • Model collapse is a documented failure mode in which a model trained on recursively generated data progressively loses the ability to produce diverse, high-quality output, with rare cases disappearing first.
  • The effect was demonstrated by Shumailov and colleagues in Nature in July 2024 and has since been examined and debated in follow-up commentary.
  • Synthetic data is safe when its origin is known and it stays anchored to a real, human-verified core; it is risky when it is unlabelled and accumulated without anyone tracking where it came from.
  • Aggregate benchmarks stay healthy through the early stages of collapse, so diversity and rare-case performance have to be measured directly.
  • Synthetic records carry no automatic privacy guarantee; NIST has stated that most synthetic-data techniques provide none unless built with differential privacy.

Is real training data always better than synthetic data?

Neither is universally better; each is suited to different conditions. Real data captures authentic complexity — sensor noise, unexpected behaviour, naturally occurring edge cases, cultural variation — and is the reference against which synthetic fidelity gets judged, but it is expensive, slow, sometimes scarce and often sensitive. Synthetic data is controllable, cheap to vary and available on demand, and it can also reproduce its generator's biases while looking entirely realistic. The question worth answering is not which is better but under what conditions generated data degrades the model that learns from it, and there is now a published answer to that.

What did the Nature study on model collapse actually show?

It showed that repeatedly training a model on its own generation's output narrows the range of what the model can produce, with rare outcomes vanishing first. In July 2024, Shumailov, Shumaylov, Zhao, Papernot, Anderson and Gal published "AI models collapse when trained on recursively generated data" in Nature (volume 631, pages 755–759). The researchers trained a model, generated data from it, trained the next generation on that output, and repeated the cycle. Across successive generations the output distribution narrowed: low-probability events — the tails — disappeared first, because they are under-represented in any finite sample and therefore under-represented in what the next generation learns from. The authors' conclusion is that indiscriminately training generative AI on a mixture of real and generated content, as happens when data is scraped from the internet, can collapse a model's ability to generate diverse, high-quality output.

Collapse is hard to notice because it begins in the tails. A model losing its ability to produce rare-but-valid outputs still performs well on average, on common cases, and on aggregate benchmarks — an argument for evaluating rare cases specifically rather than relying on averages.

The result has also been examined critically, which is normal and useful. A note published on arXiv (preprint 2410.12954, October 2024, not peer-reviewed) argues that particular assumptions in the original setup bear on how strongly the conclusion generalises. The practical position: the mechanism is real and well-motivated, its severity in any specific pipeline depends on how that pipeline curates data, and the right response is curation rather than panic or dismissal.

When is synthetic data the right choice?

Synthetic data is not a degraded substitute for real data in every application; there are cases where it is straightforwardly the better choice, and treating the collapse finding as a blanket prohibition gives those cases up for no benefit.

Use case Why synthetic works The condition
Rare events — accidents, faults, hazards Real examples are scarce, dangerous or unethical to collect Validate against the real examples that do exist; never let synthetic define the category alone
Privacy-sensitive domains Removes personal records from the training set Verify the synthesis does not leak identifiable attributes from its source
Class balancing Cheap augmentation of under-represented classes Cap the synthetic share and check whether minority-class behaviour actually improves
Simulation for perception and robotics Perfect ground truth, unlimited variation, controllable conditions Measure the sim-to-real gap explicitly and close it with real validation data
Distillation from a stronger model Deliberate transfer from a known, higher-quality teacher Not recursive self-training — the teacher outranks the student and the lineage is known
Formats and structure Templates and structured variation are cheap to generate correctly Content still needs human grounding; structure is not substance

Distillation is training a smaller or newer model on the outputs of a stronger, known teacher model, as covered in what to buy when choosing RLHF, SFT or distillation. The pattern across the table: synthetic data is safe when its lineage is known and it is anchored to something real, and dangerous when it is unlabelled, accumulated, and validated only against other synthetic data.

One assumption worth retiring here: synthetic does not automatically mean private. NIST has stated that many synthetic-data techniques provide no formal privacy guarantee, so when privacy is the objective, the generation method and its residual risk have to be evaluated rather than assumed away because the records are artificial.

Why are organisations exposed to model collapse without having chosen it?

Because ordinary practices — scraping the web, fine-tuning on prior output, using model-assisted labels — quietly build the same recursive loop the Nature study describes, even for teams that never intended to train a foundation model. The exposure is common and indirect:

  • Scraped corpora. Any dataset assembled from the open web after roughly 2023 contains machine-generated text in unknown proportion, unlabelled. Fine-tuning on it is recursive training whether or not that was the intention, a risk covered further in cleaning and deduplicating a pretraining corpus.
  • Fine-tuning on your own outputs. Teams that publish AI-assisted content and later fine-tune on their own published corpus have built a small recursive loop with no external anchor.
  • Evaluation sets contaminated by generation. Benchmarks assembled from web text may contain generated items, which flatters models producing similar text.
  • Annotation anchored to model suggestions. Model-assisted labelling accepted without unassisted control batches encodes the pre-labelling model's distribution into the dataset — the same mechanism reached from a different direction.
  • Retrieval corpora. A retrieval-augmented system whose index is full of machine-written pages retrieves machine-written answers, with no training involved at all.

The common factor is missing provenance. None of these is dangerous if it is known which items are human-authored and which are not, because then they can be weighted, filtered or excluded — a practical argument for provenance discipline entirely separate from the legal one. A related trap sits in evaluation: if the same family of models both generates the data and judges it, the system rewards its own assumptions, and the circularity is invisible from inside the scores.

How can synthetic data be used without degrading the model?

By composing the training mix deliberately with known lineage, rather than either avoiding synthetic data or accumulating it unnoticed. Seven practices make that concrete:

  1. Label every item with its origin. Human-authored, model-generated, model-assisted, or unknown. "Unknown" is a legitimate and important category — track it rather than silently treating it as human.
  2. Anchor the mix to verified human data. Keep a substantial, curated core of human-authored, human-verified material that does not shrink as synthetic volume grows. This is what stops distributional drift, and its value rises as the open web becomes less reliable as a source.
  3. Set and enforce a synthetic share. Decide the proportion deliberately, per training run and per domain, and record it. A mix nobody chose is a mix nobody can debug when behaviour changes.
  4. Filter on quality, not just validity. Generated items that are well-formed but generic drive distributional narrowing; filtering for diversity and informativeness matters more than filtering for correctness alone.
  5. Evaluate on the tails. Build evaluation sets specifically from rare cases, minority classes and edge conditions, a discipline covered in measuring dataset diversity and mining long-tail and edge cases. Aggregate metrics stay healthy through the early stages of collapse; rare-case performance does not.
  6. Keep an untouched real-data holdout. A human-authored evaluation set, collected once, never used for training, never regenerated and never augmented — the only stable reference point across model generations.
  7. Version the data, not only the model. Record which mix produced which checkpoint. When behaviour degrades, the data composition is usually the explanation, and it is unrecoverable if it was not recorded.

The synthetic share is the fraction of machine-generated items in the total training mix. The right ratio is determined experimentally for the target task, not by a universal rule, and should be reported alongside segment-level performance on the segments the synthetic data was added to improve — if those segments did not move, the augmentation did nothing except change the distribution. Compare this to how can more data make AI worse treats the broader question of when additional volume stops helping.

What does this mean for organisations publishing content at scale?

It raises the cost of publishing fluent, generic, unverified content, because that content becomes part of the corpus future models learn from. If the open web fills with machine-generated material, the corpus future models learn from degrades, and organisations producing that material are contributing to the degradation. Publishing at scale without verification is therefore not only a weak marketing asset; it is a small contribution to a collective problem.

The constructive reading raises the value of two things producers control. First, human-verified substance: original data, first-hand experience, genuine expertise, real in-market linguistic work — material that stays valuable precisely because it is not a resample of what already exists. Second, provenance marking, which lets downstream curators distinguish generated material from authored material at almost no cost at export time.

How does Lifewood approach synthetic data and model collapse?

By recording origin rather than assuming it, on both sides of the problem it works on: collecting and annotating human-authored training data, and producing AI-generated content under human direction. Human verification is the anchor in both directions, held to a 95%+ accuracy threshold under dual-layer human-in-the-loop review — reviewed twice with timestamped approval records, as detailed in AI data validation. That anchor is only producible at the network behind it: 100+ languages, 40+ delivery centres across 30+ countries and 56,000+ registered contributors, which is what it takes to generate genuinely new human material rather than recycled web text. For AI services generally, see Lifewood's AI data services.

Frequently asked questions

A degenerative process in which a generative model trained on data produced by previous generations of models progressively loses the ability to produce diverse, high-quality output. Rare events at the tails of the distribution disappear first, then the output distribution narrows overall. It was demonstrated in a 2024 *Nature* paper by Shumailov and colleagues.

No. The finding concerns indiscriminate recursive training. Synthetic data used deliberately — for rare events, privacy-sensitive domains, class balancing, simulation, or distillation from a stronger teacher — is a normal technique, provided its lineage is known, it is anchored to real human-verified data, and it is evaluated against real data.

Not from aggregate benchmarks, which stay healthy through the early stages. Watch output diversity, performance on rare classes and edge cases, and behaviour on a human-authored holdout set that has never been used for training and never regenerated. Narrowing variety at unchanged average scores is the signature.

Possibly, in two ways. If your corpus includes content your organisation produced with AI assistance, fine-tuning on it is a small recursive loop. If it includes anything scraped from the web after roughly 2023, it contains machine-generated material in unknown proportion. Neither is fatal; both are reasons to label origin and keep a human-authored anchor.

It has been examined and refined, which is how published results are supposed to be treated. A note on arXiv argues that specific assumptions in the original setup bear on how far the conclusion generalises. The mechanism is well-motivated either way, and the practical guidance — know your data's origin, anchor to human-verified material, evaluate on tails — holds under both readings.

Not automatically. The source data used to build or condition a generator may still carry rights, privacy and consent obligations, and NIST has stated that many synthetic-data techniques provide no formal privacy guarantee. Where privacy is the reason for synthesising, the method has to be evaluated on that basis rather than assumed safe.

Sources and further reading

  1. Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 631, 755–759 (2024)
  2. Borji, "A Note on Shumailov et al. (2024): AI Models Collapse When Trained on Recursively Generated Data", arXiv:2410.12954
  3. NIST, "Differentially Private Synthetic Data"

Have an AI or visibility project in mind?

From AI evaluation and human-in-the-loop review to GEO and AEO strategy, our team can help you deploy with confidence and get found in the AI search era.

Talk to our team