paper-with-me

홈 › Papers

LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models

2025-07-14 · Dachuan Shi, Yonggan Fu, Xiangchi Yuan, Zhongzhi Yu, Haoran You, Sixu Li, Xin Dong, Jan Kautz, Pavlo Molchanov, Yingyan, Lin

Recent advancements in Large Language Models (LLMs) have spurred interest in numerous applications requiring robust long-range capabilities, essential for processing extensive input contexts and continuously generating extended outputs. As sequence lengths increase, the number of Key-Value (KV) pairs in LLMs escalates, creating a significant efficiency bottleneck. In this paper, we propose a new KV cache optimization paradigm called LaCache, a training-free method for efficient and accurate generative inference of LLMs. LaCache enables LLMs to simultaneously address both of the critical challenges in long-range modeling: robust long-range capabilities and continuous generation without running out-of-memory (OOM). Specifically, LaCache integrates two key innovations: (1) a ladder-shaped KV cache pattern that stores KV pairs not only sequentially (left-to-right within each layer) but also across layers (from shallow to deep), providing an extended span for capturing long-range dependencies under a fixed storage budget, thereby boosting long-range capabilities; and (2) an iterative compaction mechanism that progressively compresses older caches, freeing up space for new tokens within a fixed cache size. This token distance-based dynamic compression enables more effective continuous generation under constrained cache budgets. Experiments across various tasks, benchmarks, and LLM models consistently validate LaCache's effectiveness in enhancing LLMs' long-range capabilities. Our code is available at https://github.com/GATECH-EIC/LaCache.

📄 PDF Abstract BibTeX arXiv:2507.14204

Code (0)

등록된 구현이 없습니다.

Tasks

Long-range modeling

Similar Papers 제목 키워드 기반

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models

2026-07-16 · Xingru Chen, Zelang Liang, Yongjia Ma, Jiqing Zhan 외 arxiv

Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However, current methods suffer from severe operator-level redundancy: they recompute the…

Text Generation

SkyLadder: Better and Faster Pretraining via Context Window Scheduling

2025-03-19 · Tongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen 외

Recent advancements in LLM pretraining have featured ever-expanding context windows to process longer sequences. However, our pilot study reveals that models pretrained with shorter context windows consistently outperfor…

8kScheduling

Efficient Ladder-style DenseNets for Semantic Segmentation of Large Images

2019-05-14 · Ivan Krešo, Josip Krapac, Siniša Šegvić

Recent progress of deep image classification models has provided great potential to improve state-of-the-art performance in related computer vision tasks. However, the transition to semantic segmentation is hampered by s…

image-classificationImage ClassificationSemantic Segmentation

LADDER: Revisiting the Cosmic Distance Ladder with Deep Learning Approaches and Exploring its Applications

2024-01-30 · Rahul Shah, Soumadeep Saha, Purba Mukherjee, Utpal Garain 외

We investigate the prospect of reconstructing the ''cosmic distance ladder'' of the Universe using a novel deep learning framework called LADDER - Learning Algorithm for Deep Distance Estimation and Reconstruction. LADDE…

Deep Learning

Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks

2026-01-09 · Elias Lumer, Faheem Nizar, Akshaya Jangiti, Kevin Frank 외 arxiv

Recent advancements in Large Language Model (LLM) agents have enabled complex multi-turn agentic tasks requiring extensive tool calling, where conversations can span dozens of API calls with increasingly large context wi…