ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
Large Language Models (LLMs) suffer significant performance degradation in multi-turn conversations when information is presented incrementally. Given that multi-turn conversations characterize everyday interactions with LLMs, this degradation poses a severe challenge to real world usability. We hypothesize that abrupt increases in model uncertainty signal misalignment in multi-turn LLM interactions, and we exploit this insight to dynamically realign conversational context. We introduce ERGO (Entropy-guided Resetting for Generation Optimization), which continuously quantifies internal uncertainty via Shannon entropy over next token distributions and triggers adaptive prompt consolidation when a sharp spike in entropy is detected. By treating uncertainty as a first class signal rather than a nuisance to eliminate, ERGO embraces variability in language and modeling, representing and responding to uncertainty. In multi-turn tasks with incrementally revealed instructions, ERGO yields a 56.6% average performance gain over standard baselines, increases aptitude (peak performance capability) by 24.7%, and decreases unreliability (variability in performance) by 35.3%, demonstrating that uncertainty aware interventions can improve both accuracy and reliability in conversational AI.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Geometric Brownian Motion under Stochastic Resetting: A Stationary yet Non-ergodic Process
We study the effects of stochastic resetting on geometric Brownian motion (GBM), a canonical stochastic multiplicative process for non-stationary and non-ergodic dynamics. Resetting is a sudden interruption of a process,…
Income inequality and mobility in geometric Brownian motion with stochastic resetting: theoretical results and empirical evidence of non-ergodicity
We explore the role of non-ergodicity in the relationship between income inequality, the extent of concentration in the income distribution, and mobility, the feasibility of an individual to change their position in the …
From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation
Combining Chain-of-Thought (CoT) with Reinforcement Learning (RL) improves text-to-image (T2I) generation, yet the underlying interaction between CoT's exploration and RL's optimization remains unclear. We present a syst…
Reinforcement LearningImage GenerationREVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
Automated chest X-ray report generation has recently benefited from reinforcement learning (RL) and large language models. However, RL training often suffers from instability or limited exploration due to fixed Kullback-…
Reinforcement LearningStochastic Resetting Accelerates Policy Convergence in Reinforcement Learning
Stochastic resetting, where a dynamical process is intermittently returned to a fixed reference state, has emerged as a powerful mechanism for optimizing first-passage properties. Existing theory largely treats static, n…
Reinforcement LearningContinuous Control