paper-with-me

홈 › Papers

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory

2026-06-23 · Habibullah Akbar arxiv

Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-value sequence grows, softmax probability mass can dilute across a wider distribution, inducing activation shift and long-context performance collapse. Moreover, long-context language modeling faces a structural tension: a sliding-window attention core maintains a bounded local representation and low perplexity but is blind to long-range dependencies, while full-context attention preserves global recall but suffers from out-of-distribution perplexity explosion. To resolve these limitations, we introduce ATMA, a hybrid convolutional-attention architecture that integrates a novel three-channel attention mechanism. ATMA factorizes the attention mixing step into: (1) a count-blind, unit-vector direction channel, (2) a bounded magnitude channel driven by the participation ratio of effective matches over an extreme-value-corrected null sink, and (3) a long-term recurrent compression memory optimized via a gated-delta fast-weights rule. Neither the Polar Attention core nor the recurrent memory is sufficient alone; their combination enables monotonic perplexity reduction and high-fidelity long-range retrieval simultaneously. We evaluate ATMA using a 120-run factorial ablation sweep, demonstrating that the combined Polar + memory model maintains induction needle-in-a-haystack retrieval accuracy above 90% out to 64K tokens (32 times the training length of 2K) while its document perplexity improves monotonically, outperforming softmax-based memory baselines which collapse at extreme context lengths. Code: https://github.com/kreasof-ai/atma

📄 PDF Abstract BibTeX arXiv:2606.25156

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity

2025-05-20 · Susav Shrestha, Brad Settlemyer, Nikoli Dryden, Narasimha Reddy

Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamically activates only a small subset of the m…

GPULarge Language Model

Nested Construction of Polar Codes via Transformers

2024-01-30 · Sravan Kumar Ankireddy, S Ashwin Hebbar, Heping Wan, Joonyoung Cho 외

Tailoring polar code construction for decoding algorithms beyond successive cancellation has remained a topic of significant interest in the field. However, despite the inherent nested structure of polar codes, the use o…

IHS-RD-Belarus at SemEval-2016 Task 5: Detecting Sentiment Polarity Using the Heatmap of Sentence

2016-06-01 · SEMEVAL 2016 6 · Maryna Chernyshevich
Aspect-Based Sentiment Analysis (ABSA)SentenceSentiment Analysis

Polarized Self-Attention: Towards High-quality Pixel-wise Regression

2021-07-02 · arXiv preprint 2021 7 · Huajun Liu, Fuqiang Liu, Xinyi Fan, Dong Huang

Pixel-wise regression is probably the most common problem in fine-grained computer vision tasks, such as estimating keypoint heatmaps and segmentation masks. These regression problems are very challenging particularly be…

2D Pose EstimationKeypoint DetectionPose Estimationregression+3

TricubeNet: 2D Kernel-Based Object Representation for Weakly-Occluded Oriented Object Detection

2021-04-23 · Beomyoung Kim, Janghyeon Lee, Sihaeng Lee, Doyeon Kim 외

We present a novel approach for oriented object detection, named TricubeNet, which localizes oriented objects using visual cues ($i.e.,$ heatmap) instead of oriented box offsets regression. We represent each object as a …

Objectobject-detectionObject DetectionObject Detection In Aerial Images+2