paper-with-me

홈 › Papers

Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs

2025-12-08 · Xiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Zhaoxiang Liu, Shiguo Lian, Ziwei He, Xipeng Qiu arxiv

Rotary Position Embeddings (RoPE) have become a standard for encoding sequence order in Large Language Models (LLMs) by applying rotations to query and key vectors in the complex plane. Standard implementations, however, utilize only the real component of the complex-valued dot product for attention score calculation. This simplification discards the imaginary component, which contains valuable phase information, leading to a potential loss of relational details crucial for modeling long-context dependencies. In this paper, we propose an extension that re-incorporates this discarded imaginary component. Our method leverages the full complex-valued representation to create a dual-component attention score. We theoretically and empirically demonstrate that this approach enhances the modeling of long-context dependencies by preserving more positional information. Furthermore, evaluations on a suite of long-context language modeling benchmarks show that our method consistently improves performance over the standard RoPE, with the benefits becoming more significant as context length increases. The code is available at https://github.com/OpenMOSS/rope_pp.

📄 PDF Abstract BibTeX arXiv:2512.07525

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Selective Rotary Position Embedding

2025-11-21 · Sajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter 외 arxiv

Position information is essential for language modeling. In softmax transformers, Rotary Position Embeddings (\textit{RoPE}) encode positions through \textit{fixed-angle} rotations, while in linear transformers, order is…

Learning to Rotate: Temporal and Semantic Rotary Encoding for Sequential Modeling

2026-04-27 · Hailing Cheng, Daqi Sun, Xinyu Lu arxiv

Every Transformer architecture dedicates enormous capacity to learning rich representations in semantic embedding space -- yet the rotation manifold acted upon by Rotary Positional Embeddings (RoPE) has been treated as a…

Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation

2025-10-12 · Jiaye Li, Baoyou Chen, Hui Li, Zilong Dong 외 arxiv

Transformers rely on explicit positional encoding to model structure in data. While Rotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-…

Text-to-Image GenerationObject Counting

MrRoPE: Mixed-radix Rotary Position Embedding

2026-01-28 · Qingyuan Tian, Wenhong Zhu, Xiaoran Liu, Xiaofeng Wang 외 arxiv

Rotary Position Embedding (RoPE)-extension refers to modifying or generalizing the Rotary Position Embedding scheme to handle longer sequences than those encountered during pre-training. However, current extension strate…

YaRN: Efficient Context Window Extension of Large Language Models

2023-08-31 · Bowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico Shippole

Rotary Position Embeddings (RoPE) have been shown to effectively encode positional information in transformer-based language models. However, these models fail to generalize past the sequence length they were trained on.…

Position