Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning
Model collapse is a central challenge in learning from synthetic data: as later-generation large language models (LLMs) are trained on an increasing proportion of model-generated data, performance can degrade due to narrowed coverage and accumulated bias. Existing work mainly studies how to bound this degradation. In iterative model evolution, however, the more meaningful objective is to ensure that each successive model improves over its predecessor, which requires diagnosing collapse at a granularity that is actionable for data curation. We study this problem in synthetic data self-improving
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%IEEE Rolls Out Large Language Models Virtual Training Course →
- FuzzyOverlapping authors or contributors · 62%pytorch/pytorch →
“Shared author/contributor keys: zou”
- FuzzyOverlapping authors or contributors · 62%mudler/LocalAI →
“Shared author/contributor keys: guo”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: luo”
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Learning from Synthetic Data without Model Collapse in Itera” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Xiaonan Luo →
“Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning”
- LinkedLinked via arxiv author · 85%Ziyue Huang →
“Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning”
- LinkedLinked via arxiv author · 85%Kehan Guo →
“Learning from Synthetic Data without Model Collapse in Iterative Instruction Tuning”
