paper-with-me

홈 › Papers

Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression

2026-03-06 · Warren Johnson arxiv

Prompt compression is often evaluated by input-token reduction, but its real deployment impact depends on how compression changes output length and total inference cost. We present a controlled replication and extension study of benchmark-dependent output dynamics under aggressive compression, covering 5,400 API calls across three benchmarks and multiple providers. To explain conflicting prior observations, we formalize instruction survival probability (Psi), a structural metric that captures whether task-critical prompt segments remain after truncation. Results show a strong benchmark effect: under r=0.3, DeepSeek exhibits severe output expansion on MBPP (56x, Psi approx 0.15) but substantially lower expansion on HumanEval (5x, Psi approx 0.72), while GPT-4o-mini is comparatively stable across benchmarks. This reconciles the apparent discrepancy between previously reported extreme explosion and lower replication effects by identifying prompt structure, not provider identity alone, as the primary moderator. We introduce the Compression Robustness Index (CRI) for cross-benchmark evaluation and show that single-benchmark assessments can produce misleading conclusions about compression safety and efficiency. To contextualize energy claims, we incorporate companion direct NVML measurements from rented RunPod GPUs and show that token savings can overstate joule savings. These findings motivate benchmark-diverse testing and structure-aware compression policies for reliable, energy-conscious LLM deployment.

📄 PDF Abstract BibTeX arXiv:2603.23527

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

2026-05-28 · Lorenz Kutschka, Bernhard Geiger arxiv

Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default language for that exchange, JSON, was designed for application-to-applicati…

Information Bottleneck: Exact Analysis of (Quantized) Neural Networks

2021-06-24 · ICLR 2022 4 · Stephan Sloth Lorenzen, Christian Igel, Mads Nielsen

The information bottleneck (IB) principle has been suggested as a way to analyze deep neural networks. The learning dynamics are studied by inspecting the mutual information (MI) between the hidden layers and the input a…

The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression

2026-03-06 · Warren Johnson arxiv

The rapid proliferation of Large Language Models has created an environmental paradox: the very technology that could help solve climate challenges is itself becoming a significant contributor to global carbon emissions.…

Basis Matters: Better Communication-Efficient Second Order Methods for Federated Learning

2021-11-02 · Xun Qian, Rustem Islamov, Mher Safaryan, Peter Richtárik

Recent advances in distributed optimization have shown that Newton-type methods with proper communication compression mechanisms can guarantee fast local rates and low communication cost compared to first order methods. …

Distributed OptimizationFederated LearningSecond-order methods

Where Matters More Than What: Decoding-aligned KV Cache Compression via Position-aware Pseudo Queries

2026-03-12 · Zhenxu Tian, Yi Su, Juntao Li, Min Zhang arxiv

The Key-Value (KV) cache is crucial for efficient Large Language Models (LLMs) inference, but excessively long contexts drastically increase KV cache memory footprint. Existing KV cache compression methods typically rely…