paper-with-me

홈 › Papers

Evaluating Zero-Shot Long-Context LLM Compression

2024-06-10 · Chenyu Wang, Yihan Wang

This study evaluates the effectiveness of zero-shot compression techniques on large language models (LLMs) under long-context. We identify the tendency for computational errors to increase under long-context when employing certain compression methods. We propose a hypothesis to explain the varied behavior of different LLM compression techniques and explore remedies to mitigate the performance decline observed in some techniques under long-context. This is a course report for COS 598D Machine Learning and Systems by Prof. Kai Li at Princeton University. Due to limited computational resources, our experiments were conducted only on LLaMA-2-7B-32K.

📄 PDF Abstract BibTeX arXiv:2406.06773

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ZeroMerge: Parameter-Free KV Cache Compression for Memory-Efficient Long-Context LLMs

2025-03-13 · Xin Liu, Pei Liu, Guoming Tang

The linear growth of key-value (KV) cache memory and quadratic computational complexity pose significant bottlenecks for large language models (LLMs) in long-context processing. While existing KV cache optimization metho…

LatentPress: Context Compression Beyond Text and Vision

2026-09-01 · Zhengze Zhou, Hejian Sang hf

Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversational histories and lon…

Text Summarization

Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning

2025-03-06 · Giulio Corallo, Orion Weller, Fabio Petroni, Paolo Papotti

Incorporating external knowledge in large language models (LLMs) enhances their utility across diverse applications, but existing methods have trade-offs. Retrieval-Augmented Generation (RAG) fetches evidence via similar…

RAGRetrieval-augmented Generation

Can LLMs Maintain Fundamental Abilities under KV Cache Compression?

2025-02-04 · Xiang Liu, Zhenheng Tang, Hong Chen, Peijie Dong 외

This paper investigates an underexplored challenge in large language models (LLMs): the impact of KV cache compression methods on LLMs' fundamental capabilities. Although existing methods achieve impressive compression r…

Arithmetic ReasoningCode GenerationLong-Context UnderstandingSensitivity+1

ZeroGVC: Zero-Shot Generative Video Compression with Autoregressive Diffusion Priors

2026-06-21 · Yixin Gao, Xiaohan Pan, Lin Liu, Xin Li 외 arxiv

Recent generative video compression methods leverage powerful generative priors to achieve perceptually pleasing reconstructions. However, most existing approaches require additional training to adapt generative models t…

Video Reconstruction