paper-with-me

홈 › Papers

When Less is More: The LLM Scaling Paradox in Context Compression

2026-02-10 · Ruishan Guo, Yibing Liu, Guoxin Ma, Yan Wang, Yueyang Zhang, Long Xia, Kecheng Chen, Zhiyuan Sun, Daiting Shi arxiv

Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabilities. However, under lossy context compression in a compressor--decoder setup, we find a \textbf{\textit{Size-Fidelity Paradox}}: increasing compressor size can lessen the faithfulness of reconstructed contexts though reconstruction error decreases. Across 27 compressor setups spanning model families, scales, and compression rates, we coin this paradox arising from two dominant factors: 1) \textit{knowledge overwriting}: larger models increasingly replace source facts with their own prior beliefs, \textit{e.g.}, `the white strawberry $\to$ the red strawberry; and 2) \textit{semantic drift}: larger models tend to paraphrase or restructure content instead of reproducing it verbatim, \textit{e.g.}, Alice hit Bob $\to$ Bob hit Alice`. Interestingly, this paradox persists across varied settings, with mid-sized compressors often outperforming larger ones in faithful recovery. By analyzing the compressed memory via embedding geometry and reconstruction determinacy, we further reveal that compressors tend to organize memory across broader semantic subspaces, yielding more ambiguous representations prone to overwriting, drift, and weakened recovery. These findings complement existing evaluations of context compression and expose a breakdown of scaling laws when the objective shifts from plausible generation to faithful preservation.

📄 PDF Abstract BibTeX arXiv:2602.09789

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge

2026-08-12 · Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi arxiv

Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer cont…

Natural Language Understanding

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size

2026-04-14 · Dikshant Kukreja, Kshitij Sah, Gautam Gupta, Avinash Anand 외 arxiv

Larger language models become simultaneously better and worse at handling contextual information -- better at ignoring false claims, worse at ignoring irrelevant tokens. We formalize this apparent paradox through the fir…

Why Less is More (Sometimes): A Theory of Data Curation

2025-11-05 · Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari-Hemmat arxiv

This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as classical scaling laws suggesting ``more is …

Mathematical Reasoning

The Scaling Paradox in Human-AI Collaboration

2026-08-01 · Anyan Qi, Mengxin Wang arxiv

The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applicati…

Paradox resolved: The allometric scaling of cancer risk across species

2020-11-22 · Christopher P. Kempes, Geoffrey B. West, John W. Pepper

Understanding the cross-species behavior of cancer is important for uncovering fundamental mechanisms of carcinogenesis, and for translating results of model systems between species. One of the most famous interspecific …