paper-with-me

Papers

Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence

2025-10-18 · Bingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng Xu arxiv

Synthetic data has been increasingly used to train frontier generative models. However, recent studies raise key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model performance, a phenomenon often coined model collapse. In this paper, we investigate ways to modify the synthetic retraining process to avoid model collapse, and even possibly help reverse the trend from collapse to improvement. Our key finding is that by injecting information through an external synthetic data verifier, whether a human or a better model, synthetic retraining will not cause model collapse. Specifically, we situate our theoretical analysis in the fundamental linear regression setting, showing that verifier-guided retraining can yield near-term improvements, but ultimately drives the parameter estimate to the verifier's "knowledge center" in the long run. Our theory further predicts that, unless the verifier is perfectly reliable, these early gains will plateau and may even reverse. Indeed, our experiments across linear regression, Variational Autoencoders (VAEs) trained on MNIST, and fining-tuning SmolLM2-135M on the XSUM task confirm these theoretical insights.

📄 PDF Abstract BibTeX arXiv:2510.16657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

2025-02-13 · Kareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong 외

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cau…

Language ModelingLanguage ModellingLarge Language Model

Escaping Mode Collapse in LLM Generation via Geometric Regulation

2026-05-01 · Xin Du, Kumiko Tanaka-Ishii arxiv

Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping to gradual loss of diversity and premature trajectory convergence. W…

Text Generation

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

2026-07-17 · Bibesh Pyakurel, M. G. Sarwar Murshed arxiv

Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated wh…

Beyond Model Collapse: Scaling Up with Synthesized Data Requires Verification

2024-06-11 · Yunzhen Feng, Elvis Dohmatob, Pu Yang, Francois Charton 외

Large Language Models (LLM) are increasingly trained on data generated by other LLM, either because generated text and images become part of the pre-training corpus, or because synthetized data is used as a replacement f…

News Summarization

Escaping Barren Plateaus in Variational Quantum Algorithms Using Negative Learning Rate in Quantum Internet of Things

2025-11-28 · Ratun Rahman, Dinh C. Nguyen arxiv

Variational Quantum Algorithms (VQAs) are becoming the primary computational primitive for next-generation quantum computers, particularly those embedded as resource-constrained accelerators in the emerging Quantum Inter…