paper-with-me

홈 › Papers

Rotary Positional Embeddings as Phase Modulation: Theoretical Bounds on the RoPE Base for Long-Context Transformers

2026-02-11 · Feilong Liu arxiv

Rotary positional embeddings (RoPE) are widely used in large language models to encode token positions through multiplicative rotations, yet their behavior at long context lengths remains poorly characterized. In this work, we reinterpret RoPE as phase modulation applied to a bank of complex oscillators, enabling analysis through classical signal processing theory. Under this formulation, we derive principled lower bounds on the RoPE base parameter that are necessary to preserve positional coherence over a target context length. These include a fundamental aliasing bound, analogous to a Nyquist limit, and a DC-component stability bound that constrains phase drift in low-frequency positional modes. We further extend this analysis to deep transformers, showing that repeated rotary modulation across layers compounds angular misalignment, tightening the base requirement as depth increases. Complementing these results, we derive a precision-dependent upper bound on the RoPE base arising from finite floating-point resolution. Beyond this limit, incremental phase updates become numerically indistinguishable, leading to positional erasure even in the absence of aliasing. Together, the lower and upper bounds define a precision- and depth-dependent feasibility region a Goldilocks zone for long-context transformers. We validate the framework through a comprehensive case study of state-of-the-art models, including LLaMA, Mistral, and DeepSeek variants, showing that observed successes, failures, and community retrofits align closely with the predicted bounds. Notably, models that violate the stability bound exhibit attention collapse and long-range degradation, while attempts to scale beyond one million tokens encounter a hard precision wall independent of architecture or training.

📄 PDF Abstract BibTeX arXiv:2602.10959

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Context-aware Rotary Position Embedding

2025-07-30 · Ali Veisi, Delaram Fartoot, Hamidreza Amirzadeh arxiv

Positional encoding is a vital component of Transformer architectures, enabling models to incorporate sequence order into self-attention mechanisms. Rotary Positional Embeddings (RoPE) have become a widely adopted soluti…

Computational Efficiency

Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings

2025-12-13 · Yoav Gelberg, Koshi Eguchi, Takuya Akiba, Edoardo Cetin arxiv

So far, expensive finetuning beyond the pretraining sequence length has been a requirement for effectively extending the context of language models (LM). In this work, we break this key bottleneck by Dropping the Positio…

Clifford Algebraic Rotor Embeddings : Maybe embeddings should start to CARE

2025-11-11 · Sameeksha Sriram, Ayush Paliwal, Alexander S. Ecker, Chase van de Geijn arxiv

Rotary Positional Embeddings (RoPE) have demonstrated exceptional performance as a positional encoding method, consistently outperforming their baselines. While recent work has sought to extend RoPE to higher-dimensional…

Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs

2025-12-08 · Xiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang 외 arxiv

Rotary Position Embeddings (RoPE) have become a standard for encoding sequence order in Large Language Models (LLMs) by applying rotations to query and key vectors in the complex plane. Standard implementations, however,…

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance

2026-07-28 · Abraham Chachamovits arxiv

Transformer language models are usually analyzed through vector geometry, yet ordered context and rotary position encoding introduce explicit phase structure into query-key interactions. This paper develops a bounded spe…