paper-with-me

홈 › Papers

Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language Models

2024-07-22 · Alliot Nagle, Adway Girish, Marco Bondaschi, Michael Gastpar, Ashok Vardhan Makkuva, Hyeji Kim

We formalize the problem of prompt compression for large language models (LLMs) and present a framework to unify token-level prompt compression methods which create hard prompts for black-box models. We derive the distortion-rate function for this setup as a linear program, and provide an efficient algorithm to compute this fundamental limit via the dual of the linear program. Using the distortion-rate function as the baseline, we study the performance of existing compression schemes on a synthetic dataset consisting of prompts generated from a Markov chain, natural language queries, and their respective answers. Our empirical analysis demonstrates the criticality of query-aware prompt compression, where the compressor has knowledge of the downstream task/query for the black-box LLM. We show that there is a large gap between the performance of current prompt compression methods and the optimal strategy, and propose Adaptive QuerySelect, a query-aware, variable-rate adaptation of a prior work to close the gap. We extend our experiments to a small natural language dataset to further confirm our findings on our synthetic dataset.

📄 PDF Abstract BibTeX arXiv:2407.15504

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Queries

Similar Papers 제목 키워드 기반

Fundamental Limits of Communication Efficiency for Model Aggregation in Distributed Learning: A Rate-Distortion Approach

2022-06-28 · Naifu Zhang, Meixia Tao, Jia Wang, Fan Xu

One of the main focuses in distributed learning is communication efficiency, since model aggregation at each round of training can consist of millions to billions of parameters. Several model compression methods, such as…

Model CompressionQuantization

Neural Estimation of the Rate-Distortion Function With Applications to Operational Source Coding

2022-04-04 · Eric Lei, Hamed Hassani, Shirin Saeedi Bidokhti

A fundamental question in designing lossy data compression schemes is how well one can do in comparison with the rate-distortion function, which describes the known theoretical limits of lossy compression. Motivated by t…

Data Compression

Training-Free Rate-Distortion-Perception Traversal With Diffusion

2026-03-04 · Yuhan Wang, Suzhi Bi, Ying-Jun Angela Zhang arxiv

The rate-distortion-perception (RDP) tradeoff characterizes the fundamental limits of lossy compression by jointly considering bitrate, reconstruction fidelity, and perceptual quality. While recent neural compression met…

An Information-Theoretic Justification for Model Pruning

2021-02-16 · Berivan Isik, Tsachy Weissman, Albert No

We study the neural network (NN) compression problem, viewing the tension between the compression ratio and NN performance through the lens of rate-distortion theory. We choose a distortion metric that reflects the effec…

Data CompressionmodelModel Compression

Rate Distortion For Model Compression: From Theory To Practice

2018-10-09 · Weihao Gao, Yu-Han Liu, Chong Wang, Sewoong Oh

The enormous size of modern deep neural networks makes it challenging to deploy those models in memory and communication limited scenarios. Thus, compressing a trained model without a significant loss in performance has …

Data CompressionModel CompressionQuantization