Research
Studies find synthetic data safe at scale — if curated like the real thing
New research overturns the bleakest 'model collapse' predictions, showing recursively trained models stay healthy when synthetic data is filtered and mixed carefully.
By Priya Sharma, Research Editor — MONTREAL
MONTREAL — The feared death spiral of models trained on model output — 'model collapse' — looks avoidable after all. Two large-scale studies published this week find that recursive training remains stable across many generations, provided synthetic data is quality-filtered and blended with fresh real data rather than replacing it.
The results matter commercially as much as scientifically: high-quality human text is increasingly licensed, expensive or exhausted, and every major lab already leans on synthetic corpora for reasoning and code.
Both papers caution that diversity, not accuracy, is the quantity to guard: models fed only their own consensus drift subtly toward blandness long before any measurable benchmark decline.
Enable JavaScript to read the full story on Neural Daily News.