Knowing When to Stop: Dynamic Context Cutoff for Large Language Models
Large language models (LLMs) process entire input contexts indiscriminately, which is inefficient in cases where the information required to answer a query is localized within the context. We present dynamic context cutoff, a human-inspired method enabling LLMs to self-terminate processing upon acquiring sufficient task-relevant information. Through analysis of model internals, we discover that specific attention heads inherently encode "sufficiency signals" - detectable through lightweight classifiers - that predict when critical information has been processed. This reveals a new efficiency paradigm: models' internal understanding naturally dictates processing needs rather than external compression heuristics. Comprehensive experiments across six QA datasets (up to 40K tokens) with three model families (LLaMA/Qwen/Mistral, 1B0-70B) demonstrate 1.33x average token reduction while improving accuracy by 1.3%. Furthermore, our method demonstrates better performance with the same rate of token reduction compared to other context efficiency methods. Additionally, we observe an emergent scaling phenomenon: while smaller models require require probing for sufficiency detection, larger models exhibit intrinsic self-assessment capabilities through prompting.
Code (0)
등록된 구현이 없습니다.
Tasks
Token ReductionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
We investigate whether the success of a zero-shot Chain-of-Thought (CoT) process can be predicted before completion. We discover that a probing classifier, based on LLM representations, performs well \emph{even before a …
Explorative Data Analysis of Time Series based AlgorithmFeatures of CMA-ES Variants
In this study, we analyze behaviours of the well-known CMA-ES by extracting the time-series features on its dynamic strategy parameters. An extensive experiment was conducted on twelve CMA-ES variants and 24 test problem…
Clusteringfeature selectionTime SeriesTime Series AnalysisLow Profile Metamaterial Band-Pass Filter Loaded with 4-Turn Complementary Spiral Resonator for WPT Applications
In this paper, a very compact and low insertion loss metamaterial band-pass filter (MBPF) at the center frequency of f0=730 MHz is proposed, based on the rectangular-shape 4 turn complementary spiral resonators (4 CSR). …
Turning the Ratchet: Dynamic Screening with Multiple Agents
We study a dynamic contracting problem with multiple agents and limited commitment. A principal seeks to screen efficient agents using one-period contracts, but is tempted to revise contract terms upon knowing an agent's…
Smooth Dynamic Cutoffs for Machine Learning Interatomic Potentials
Machine learning interatomic potentials (MLIPs) have proven to be wildly useful for molecular dynamics simulations, powering countless drug and materials discovery applications. However, MLIPs face two primary bottleneck…