paper-with-me

홈 › Papers

Does the LM Head Create a Harmful Gradient Bottleneck? A Causal Test

2026-08-17 · Anand Murugan arxiv

The language-model head maps a hidden state of width D to a vocabulary of size V, so its transpose can return at most D independent directions to the Transformer. Godey and Artzi argue that this severe projection is a harmful optimization bottleneck. We separate the geometry from the causal claim. Our backward-only intervention keeps the ordinary logits and the exact LM-head parameter update while reducing only the rank of the gradient sent into the Transformer. Across five paired seeds on byte-level and BPE-8192 WikiText-2 models, reducing backward rank increases validation loss. An equally ranked factorized forward head, however, increases loss substantially more. At half rank in the larger model, the backward-only loss increase is 0.0586 (95% CI [0.0167, 0.1005]), while the factorized forward head increases loss by 0.1795 ([0.1547, 0.2042]). The vocabulary-space residual also contributes to the ordinary LM-head update, and removing that contribution is harmful. Additional controls show that repeated-token failures are confounded by the number of independently sampled symbols, that adding never-target output classes does not impair learning, and that projection diagnostics do not reliably predict progress in our runs. Tested auxiliary feedback routes do not beat tuned backpropagation. These results confirm strong geometric compression but do not establish that it is a harmful optimization bottleneck.

📄 PDF Abstract BibTeX arXiv:2608.16671

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

2026-05-19 · Yuchun Miao, Sen Zhang, Yuqi Zhang, Yaorui Shi 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rollout samples are expensive to obtain, making sample efficiency a criti…

Reinforcement Learning

Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks

2026-05-07 · Guoxin Lu, Letian Sha, Qing Wang, Peijie Sun 외 arxiv

The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Fine-tuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they…

Lost in Backpropagation: The LM Head is a Gradient Bottleneck

2026-03-10 · Nathan Godey, Yoav Artzi arxiv

The last layer of neural language models (LMs) projects output features of dimension $D$ to logits in dimension $V$, the size of the vocabulary, where usually $D \ll V$. This mismatch is known to raise risks of limited e…

Self-Destructive Language Model

2025-05-18 · Yuhui Wang, Rongyi Zhu, Ting Wang

Harmful fine-tuning attacks pose a major threat to the security of large language models (LLMs), allowing adversaries to compromise safety guardrails with minimal harmful data. While existing defenses attempt to reinforc…

Language ModelingLanguage Modellingmodel

Quantizing data for distributed learning

2020-12-14 · Osama A. Hanna, Yahya H. Ezzeldin, Christina Fragouli, Suhas Diggavi

We consider machine learning applications that train a model by leveraging data distributed over a trusted network, where communication constraints can create a performance bottleneck. A number of recent approaches propo…

Quantization