paper-with-me

홈 › Papers

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

2026-08-04 · Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer hf

We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes out a large fraction of attention weights and renders the affected attention heads partially blind. We analyze this failure mode, characterize its impact, and examine four mitigation strategies. We further demonstrate its occurrence in state-of-the-art pretrained models based on ALiBi. Comprehensive pretraining experiments with 148M-parameter decoder models help us to disentangle its effects from out-of-context degradation. We find that ALiBi's failure mode can substantially impair token retrieval while having only a minor effect on standard decoder benchmarks. We propose four training-time mitigation strategies and evaluate them individually and in combinations, finding that log-scaled distances yield the most consistent improvements in passkey retrieval. Despite this problem, default ALiBi slopes remain a surprisingly strong baseline, particularly for needle-in-a-haystack retrieval. Based on these findings we provide concrete recommendations on how to train models with ALiBi.

📄 PDF Abstract BibTeX arXiv:2608.03994

Code (1)

Tavish9/awesome-daily-AI-arxiv ★ 113

Similar Papers 제목 키워드 기반

Breaking Spatial Boundaries: Spectral-Domain Registration Guided Hyperspectral and Multispectral Blind Fusion

2025-06-25 · Kunjing Yang, Libin Zheng, Minru Bai, Ting Lu 외

The blind fusion of unregistered hyperspectral images (HSIs) and multispectral images (MSIs) has attracted growing attention recently. To address the registration challenge, most existing methods employ spatial transform…

Discovering and Validating AI Errors With Crowdsourced Failure Reports

2021-09-23 · Ángel Alexander Cabrera, Abraham J. Druck, Jason I. Hong, Adam Perer

AI systems can fail to learn important behaviors, leading to real-world issues like safety concerns and biases. Discovering these systematic failures often requires significant developer attention, from hypothesizing pot…

Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning

2026-04-20 · Moiz Imran, Sahan Bulathwela arxiv

Intelligent tutoring systems increasingly provide automated feedback on student work, but robust feedback requires assessing reasoning, not only final answers. We study a failure mode we call the correct answer trap (CAT…

Failure Ontology: A Lifelong Learning Framework for Blind Spot Detection and Resilience Design

2026-04-12 · Yuan Sun, Hong Yi, Jinyuan Liu arxiv

Personalized learning systems are almost universally designed around a single objective: help people acquire knowledge and skills more efficiently. We argue this framing misses the more consequential problem. The most da…

Blind Diagnosis for Millimeter-wave Large-scale Antenna Systems

2021-01-25 · Rui Sun, Weidong Wang, Li Chen, Guo Wei 외

Millimeter-wave (mmWave) communication systems rely on large-scale antenna arrays to combat large path-loss at mmWave band. Due to hardware characteristics and deployment environments, mmWave large-scale antenna systems …

Diagnostic