paper-with-me

홈 › Papers

COMI: Coarse-to-fine Context Compression via Marginal Information Gain

2026-02-02 · Jiwei Tang, Shilei Liu, Zhicheng Zhang, Yujin Yuan, Libin Zheng, Wenbo Su, Bo Zheng arxiv

Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse tasks. However, their deployment in long context scenarios remains hindered by computational inefficiency and information redundancy. Context compression methods address these challenges by significantly reducing input length and eliminating redundancy. We propose COMI, a coarse-to-fine adaptive context compression framework that jointly optimizes for semantic relevance and diversity under high compression rates. We introduce Marginal Information Gain (MIG), a metric defined as the relevance of a unit to the input query minus its semantic redundancy with other units, guiding the compression process to prioritize information that is both relevant and low redundant. The framework operates in two stages: (1) Coarse-Grained Group Reallocation, where the context is partitioned into groups and dynamically assigned compression rates based on inter-group MIG, ensuring compression budgets align with information value distribution; and (2) Fine-Grained Token Merging, where tokens within each group are fused via an intra-group MIG-based weighting mechanism, thereby preserving key semantics while avoiding the accumulation of redundancy. Extensive experiments across question-answering (e.g., NaturalQuestions, 2WikiMQA, HotpotQA and NarrativeQA), summarization (e.g., MultiNews) with various backbones (e.g., LLaMA-2-7B, Qwen2-7B) show that COMI outperforms existing baselines by a large margin, e.g., approximately 25-point Exact Match (EM) improvement under 32x compression constraint with Qwen2-7B on NaturalQuestions.

📄 PDF Abstract BibTeX arXiv:2602.01719

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Level Causal Embeddings

2026-02-25 · Willem Schooltink, Fabio Massimo Zennaro arxiv

Abstractions of causal models allow for the coarsening of models such that relations of cause and effect are preserved. Whereas abstractions focus on the relation between two models, in this paper we study a framework fo…

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

2023-10-09 · Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang 외

Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the …

GSM8KIn-Context Learning

LongCodeZip: Compress Long Context for Code Language Models

2025-10-01 · Yuling Shi, Yichun Qian, Hongyu Zhang, Beijun Shen 외 arxiv

Code generation under long contexts is becoming increasingly critical as Large Language Models (LLMs) are required to reason over extensive information in the codebase. While recent advances enable code LLMs to process l…

Question AnsweringCode GenerationCode Completion

Deep Lossless Image Compression via Masked Sampling and Coarse-to-Fine Auto-Regression

2025-03-14 · Tiantian Li, Qunbing Xia, Yue Li, Ruixiao Guo 외

Learning-based lossless image compression employs pixel-based or subimage-based auto-regression for probability estimation, which achieves desirable performances. However, the existing works only consider context depende…

Image Compressionregression

Context-Based Trit-Plane Coding for Progressive Image Compression

2023-03-10 · CVPR 2023 1 · Seungmin Jeon, Kwang Pyo Choi, Youngo Park, Chang-Su Kim

Trit-plane coding enables deep progressive image compression, but it cannot use autoregressive context models. In this paper, we propose the context-based trit-plane coding (CTC) algorithm to achieve progressive compress…

DecoderImage Compression