SALT: Salience-Aware Lexical Trie for Long-Context Compression
As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-level prompt compression methods address this, but rank each sentence by a scalar relevance score, treating the document as an unstructured pool of words and sentences. Under tight budgets, this causes theme collapse, where the dominant theme(s) of a document consumes the budget, discarding less-frequent yet task-relevant themes. Preserving thematic coverage instead requires allocating the budget across recurring themes rather than scoring sentences in isolation. To this end, we propose SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure. This trie-based organization smooths memory allocation and prevents dominant themes from monopolizing the budget. Multi-anchor retrieval activates trie nodes labeled by query keywords at any depth, and the trie persists across dialogue turns, supporting multi-turn use without re-encoding the document. By preserving document themes, SALT reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that target decoding-time latency and memory.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A multiple attribute model resolves a conflict between additive and multiplicative models of incentive salience
A model of incentive salience as a function of stimulus value and interoceptive state has been previously proposed. In that model, the function differs depending on whether the stimulus is appetitive or aversive; it is m…
AttributeInformation Retrieval in long documents: Word clustering approach for improving Semantics
In this paper, we propose an alternative to deep neural networks for semantic information retrieval for the case of long documents. This new approach exploiting clustering techniques to take into account the meaning of w…
ClusteringInformation RetrievalRetrievalEnergy Storage Autonomy in Renewable Energy Systems Through Hydrogen Salt Caverns
The expansion of renewable energy sources leads to volatility in electricity generation within energy systems. Subsurface storage of hydrogen in salt caverns can play an important role in long-term energy storage, but th…
Memory and Knowledge Augmented Language Models for Inferring Salience in Long-Form Stories
Measuring event salience is essential in the understanding of stories. This paper takes a recent unsupervised method for salience detection derived from Barthes Cardinal Functions and theories of surprise and applies it …
FormLanguage ModelingLanguage ModellingRetrieval+1Mitigating Posterior Salience Attenuation in Long-Context LLMs with Positional Contrastive Decoding
While Large Language Models (LLMs) support long contexts, they struggle with performance degradation within the context window. Current solutions incur prohibitive training costs, leaving statistical behaviors and cost-e…