paper-with-me

홈 › Papers

The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression

2026-03-06 · Warren Johnson arxiv

The rapid proliferation of Large Language Models has created an environmental paradox: the very technology that could help solve climate challenges is itself becoming a significant contributor to global carbon emissions. We test whether prompt compression improves inference energy efficiency in 28,421 successful API trials (28,428 planned) across three providers (OpenAI GPT-4o-mini, Anthropic Claude-3.5-Sonnet, and DeepSeek-Chat), five benchmarks (HumanEval, MBPP, GSM8K, MATH, MMLU), and four compression ratios (r in {1.0, 0.7, 0.5, 0.3}). Energy is estimated with a token-based proxy calibrated against local direct measurements, and quality is tracked with benchmark pass rates. Compression produced substantial quality loss (overall pass rate 26.0% at baseline vs. 1.5% at r=0.7) and strongly provider-dependent energy behavior. DeepSeek exhibited output expansion under compression (21 to 798 tokens at r=0.3), corresponding to energy increases up to +2,140%, while GPT-4o-mini showed mixed effects including a reduction at r=0.5. These results indicate that input-token reduction alone is not a reliable energy optimization strategy in production inference. For the evaluated settings, model selection and output-length control provided more consistent energy-quality tradeoffs than prompt compression.

📄 PDF Abstract BibTeX arXiv:2603.23528

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression

2026-03-06 · Warren Johnson arxiv

Prompt compression is often evaluated by input-token reduction, but its real deployment impact depends on how compression changes output length and total inference cost. We present a controlled replication and extension …

From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers

2025-11-19 · Huiyuan Tian, Bonan Xu, Shijian Li arxiv

Feature-map knowledge distillation (KD) transfers internal representations well between comparably sized Vision Transformers (ViTs), but it often fails in compression. We revisit this failure and uncover a paradox. Sampl…

Knowledge Distillation

Energy consumption of code small language models serving with runtime engines and execution providers

2024-12-19 · Francisco Durán, Matias Martinez, Patricia Lago, Silverio Martínez-Fernández

Background. The rapid growth of Language Models (LMs), particularly in code generation, requires substantial computational resources, raising concerns about energy consumption and environmental impact. Optimizing LMs inf…

Code GenerationCPU

Bandwidth-efficient Inference for Neural Image Compression

2023-09-06 · Shanzhi Yin, Tongda Xu, Yongsheng Liang, Yuanyuan Wang 외

With neural networks growing deeper and feature maps growing larger, limited communication bandwidth with external memory (or DRAM) and power constraints become a bottleneck in implementing network inference on mobile an…

Data CompressionImage CompressionQuantization

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

2026-05-28 · Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun, Fnu Suya arxiv

Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affects what users pay. We show that this kind of billing is hard to audit…