paper-with-me

Papers

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

2026-05-28 · Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Kyle Lampinen, Christopher Potts, Ekdeep Singh Lubana arxiv

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger model will be able to learn a part of the data distribution that a smaller model fails to learn, even with infinite training data. To validate this claim and identify its causes, we study the effects of model scaling on a synthetic setup consisting of a mixture of tasks that show monotonic scaling curves. The results point to a data-induced competition over resources (neurons). Specifically, smaller models allocate their neurons to high frequency or low complexity tasks, and so they learn solutions that perform poorly on rare and complex tasks. Moreover, this happens even when solutions capable of expressing the desired task exist. We then assess how a larger model circumvents this data-centric bottleneck, finding that it traces to a reduced interference mechanism: larger models can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate. Finally, to further validate these claims, we pretrain OLMo models (4M to 4B parameters) on novel tasks of varying frequency and complexity. The results mirror those from our synthetic data experiments: only the larger OLMo models learn the infrequent and complex tasks, and these larger models embed more task features in their representations and show less gradient interference between tasks. Overall, we offer a data-centric account of why larger models learn tasks that smaller models fail to. This helps explain why larger models are better in practice, and it can inform practical questions concerning model sizing and training data mixtures.

📄 PDF Abstract BibTeX arXiv:2605.29548

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Causal Inference under Networked Interference and Intervention Policy Enhancement

2020-02-20 · Yunpu Ma, Volker Tresp

Estimating individual treatment effects from data of randomized experiments is a critical task in causal inference. The Stable Unit Treatment Value Assumption (SUTVA) is usually made in causal inference. However, interfe…

Causal Inference

Impact of Synchronization Offsets and CSI Feedback Delay in Distributed MIMO Systems

2025-03-31 · Kumar Sai Bondada, Daniel Jakubisin, R. Michael Buehrer

The main challenges of distributed MIMO systems lie in achieving highly accurate synchronization and ensuring the availability of accurate channel state information (CSI) at distributed nodes. This paper analytically exa…

Scaling End-to-End Models for Large-Scale Multilingual ASR

2021-04-30 · Bo Li, Ruoming Pang, Tara N. Sainath, Anmol Gulati 외

Building ASR models across many languages is a challenging multi-task learning problem due to large variations and heavily unbalanced data. Existing work has shown positive transfer from high resource to low resource lan…

Multi-Task Learning

Interference Analysis for Coexistence of UAVs and Civil Aircrafts Based on Automatic Dependent Surveillance-Broadcast

2024-06-12 · Yiyang Liao, Ziye Jia, Chao Dong, Lei Zhang 외

Due to the advantages of high mobility and easy deployment, unmanned aerial vehicles (UAVs) are widely applied in both military and civilian fields. In order to strengthen the flight surveillance of UAVs and guarantee th…

A Language Model with Limited Memory Capacity Captures Interference in Human Sentence Processing

2023-10-24 · William Timkey, Tal Linzen

Two of the central factors believed to underpin human sentence processing difficulty are expectations and retrieval from working memory. A recent attempt to create a unified cognitive model integrating these two factors …

Language ModelingLanguage ModellingRetrievalSentence