paper-with-me

홈 › Papers

Latent-Condensed Transformer for Efficient Long Context Modeling

2026-04-14 · Zeng You, Yaofo Chen, Qiuwu Chen, Ying Sun, Shuhai Zhang, Yingjian Li, Yaowei Wang, Mingkui Tan arxiv

Large language models (LLMs) face significant challenges in processing long contexts due to the linear growth of the key-value (KV) cache and quadratic complexity of self-attention. Existing approaches address these bottlenecks separately: Multi-head Latent Attention (MLA) reduces the KV cache by projecting tokens into a low-dimensional latent space, while sparse attention reduces computation. However, sparse methods cannot operate natively on MLA's compressed latent structure, missing opportunities for joint optimization. In this paper, we propose Latent-Condensed Attention (LCA), which directly condenses context within MLA's latent space, where the representation is disentangled into semantic latent vectors and positional keys. LCA separately aggregates semantic vectors via query-aware pooling and preserves positional keys via anchor selection. This approach jointly reduces both computational cost and KV cache without adding parameters. Beyond MLA, LCA's design is architecture-agnostic and readily extends to other attention mechanisms such as GQA. Theoretically, we prove a length-independent error bound. Experiments show LCA achieves up to 2.5$\times$ prefilling speedup and 90% KV cache reduction at 128K context while maintaining competitive performance.

📄 PDF Abstract BibTeX arXiv:2604.12452

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling

2025-03-06 · Zhuo Chen, Oriol Mayné i Comas, Zhuotao Jin, Di Luo 외

We rigorously establish a bipartite mutual information scaling law in natural language that governs long-range dependencies. This scaling law, which we show is distinct from and scales independently of the conventional t…

Language ModelingLanguage ModellingState Space Models

Text Similarity in Vector Space Models: A Comparative Study

2018-09-24 · Shahmirzadi Omid, Lugowski Adam, Younge Kenneth

Automatic measurement of semantic text similarity is an important task in natural language processing. In this paper, we evaluate the performance of different vector space models to perform this task. We address the real…

text similarityTopic Models

Transformer-based Context Condensation for Boosting Feature Pyramids in Object Detection

2022-07-14 · Zhe Chen, Jing Zhang, Yufei Xu, DaCheng Tao

Current object detectors typically have a feature pyramid (FP) module for multi-level feature fusion (MFF) which aims to mitigate the gap between features from different levels and form a comprehensive object representat…

object-detectionObject Detection

ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling

2026-07-06 · Archie Chaudhury arxiv

Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more efficient, parallelized training across a wide set of documents and corpora. This has a…

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling

2026-04-27 · Parsa Ashrafi Fashi, Utkarsh Saxena, Mehdi Rezagholizadeh, Aref Jafari 외 arxiv

Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transformers, but most are still pretrained from scratch and therefore fail to…

Common Sense Reasoning