paper-with-me

홈 › Papers

Positional Biases Shift as Inputs Approach Context Window Limits

2025-08-10 · Blerta Veseli, Julian Chibane, Mariya Toneva, Alexander Koller arxiv

Large Language Models (LLMs) often struggle to use information across long inputs effectively. Prior work has identified positional biases, such as the Lost in the Middle (LiM) effect, where models perform better when information appears at the beginning (primacy bias) or end (recency bias) of the input, rather than in the middle. However, long-context studies have not consistently replicated these effects, raising questions about their intensity and the conditions under which they manifest. To address this, we conducted a comprehensive analysis using relative rather than absolute input lengths, defined with respect to each model's context window. Our findings reveal that the LiM effect is strongest when inputs occupy up to 50% of a model's context window. Beyond that, the primacy bias weakens, while recency bias remains relatively stable. This effectively eliminates the LiM effect; instead, we observe a distance-based bias, where model performance is better when relevant information is closer to the end of the input. Furthermore, our results suggest that successful retrieval is a prerequisite for reasoning in LLMs, and that the observed positional biases in reasoning are largely inherited from retrieval. These insights have implications for long-context tasks, the design of future LLM benchmarks, and evaluation methodologies for LLMs handling extended inputs.

📄 PDF Abstract BibTeX arXiv:2508.07479

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training

2023-09-19 · Dawei Zhu, Nan Yang, Liang Wang, YiFan Song 외

Large Language Models (LLMs) are trained with a pre-defined context length, restricting their use in scenarios requiring long inputs. Previous efforts for adapting LLMs to a longer length usually requires fine-tuning wit…

2kPosition

Positional Fragility in LLMs: How Offset Effects Reshape Our Understanding of Memorization Risks

2025-05-19 · YiXuan Xu, Antoine Bosselut, Imanol Schlag

Large language models are known to memorize parts of their training data, posing risk of copyright violations. To systematically examine this risk, we pretrain language models (1B/3B/8B) from scratch on 83B tokens, mixin…

AttributeMemorization

Exploring Context Window of Large Language Models via Decomposed Positional Vectors

2024-05-28 · Zican Dong, Junyi Li, Xin Men, Wayne Xin Zhao 외

Transformer-based large language models (LLMs) typically have a limited context window, resulting in significant performance degradation when processing text beyond the length of the context window. Extensive studies hav…

Long-Context Language Modeling with Parallel Context Encoding

2024-02-26 · Howard Yen, Tianyu Gao, Danqi Chen

Extending large language models (LLMs) to process longer inputs is crucial for a wide range of applications. However, the substantial computational cost of transformers and limited generalization of positional encoding r…

In-Context LearningInstruction FollowingLanguage ModelingLanguage Modelling

Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

2026-05-10 · Daniel Wolfson, Tal Wagner arxiv

Positional encoding in transformers is commonly implemented through positional embeddings, attention masks, or bias terms, but formal connections between these mechanisms remain limited. We study attention with positiona…