paper-with-me

홈 › Papers

SALT: Salience-Aware Lexical Trie for Long-Context Compression

2026-07-20 · Oteo Mamo, Hyunjin Yi, Joydhriti Choudhury, Shangqian Gao, Weikuan Yu arxiv

As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. Existing input-level prompt compression methods address this, but rank each sentence by a scalar relevance score, treating the document as an unstructured pool of words and sentences. Under tight budgets, this causes theme collapse, where the dominant theme(s) of a document consumes the budget, discarding less-frequent yet task-relevant themes. Preserving thematic coverage instead requires allocating the budget across recurring themes rather than scoring sentences in isolation. To this end, we propose SALT, a model-agnostic extractive framework that organizes per-sentence keywords into a trie ordered by sentence frequency (SF), a lightweight, reusable proxy for document thematic structure. This trie-based organization smooths memory allocation and prevents dominant themes from monopolizing the budget. Multi-anchor retrieval activates trie nodes labeled by query keywords at any depth, and the trie persists across dialogue turns, supporting multi-turn use without re-encoding the document. By preserving document themes, SALT reduces the prefill computation and memory cost of long-context prompts while remaining composable with KV-cache methods that target decoding-time latency and memory.

📄 PDF Abstract BibTeX arXiv:2607.17486

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A multiple attribute model resolves a conflict between additive and multiplicative models of incentive salience

2018-12-20

A model of incentive salience as a function of stimulus value and interoceptive state has been previously proposed. In that model, the function differs depending on whether the stimulus is appetitive or aversive; it is m…

Attribute

Information Retrieval in long documents: Word clustering approach for improving Semantics

2023-02-20 · Paul Mbate Mekontchou, Armel Fotsoh, Bernabe Batchakui, Eddy Ella

In this paper, we propose an alternative to deep neural networks for semantic information retrieval for the case of long documents. This new approach exploiting clustering techniques to take into account the meaning of w…

ClusteringInformation RetrievalRetrieval

Energy Storage Autonomy in Renewable Energy Systems Through Hydrogen Salt Caverns

2025-04-16 · David Franzmann, Thora Schubert, Heidi Heinrichs, Peter A. Kukla 외

The expansion of renewable energy sources leads to volatility in electricity generation within energy systems. Subsurface storage of hydrogen in salt caverns can play an important role in long-term energy storage, but th…

Memory and Knowledge Augmented Language Models for Inferring Salience in Long-Form Stories

2021-09-08 · EMNLP 2021 11 · David Wilmot, Frank Keller

Measuring event salience is essential in the understanding of stories. This paper takes a recent unsupervised method for salience detection derived from Barthes Cardinal Functions and theories of surprise and applies it …

FormLanguage ModelingLanguage ModellingRetrieval+1

Mitigating Posterior Salience Attenuation in Long-Context LLMs with Positional Contrastive Decoding

2025-06-10 · Zikai Xiao, Ziyang Wang, Wen Ma, Yan Zhang 외

While Large Language Models (LLMs) support long contexts, they struggle with performance degradation within the context window. Current solutions incur prohibitive training costs, leaving statistical behaviors and cost-e…