paper-with-me

홈 › Papers

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

2026-05-26 · Haoyu Zheng, Yun Zhu, Shu Yuan, Shangming Chen, Qing Wang, Wenqiao Zhang, Jun Xiao, Yueting Zhuang arxiv

Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lost-in-conversation (LiC) gap. We trace part of this degradation to self-contamination: intermediate assistant replies enter later context and carry early deviations forward. Motivated by this mechanism, we propose MAIGO, an on-policy self-distillation method that reduces this contamination using history-cleaned references from the model's own policy. For middle turns, MAIGO removes prior assistant replies while preserving the user-visible sharded prefix; for answer turns, it distills from paired full-view references conditioned on the completed user-side dialogue. A reliability weight downweights middle-turn samples that disagree with the clean reference. MAIGO requires no verifier rewards, state labels, or inference-time scaffolding. Under the LiC paired-view protocol with deterministic verifiers, MAIGO improves Qwen2.5-7B-Instruct SHARDED accuracy from 52.8 to 66.1 and the SHARDED/FULL ratio from 66.5% to 84.1%, while keeping FULL accuracy within 2.3 points. These results show that self-contamination is a trainable component of the LiC gap.

📄 PDF Abstract BibTeX arXiv:2605.27186

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability

2026-05-26 · Ramakrishna Vamsi Setti, Jagadeesh Rachapudi, Sachin Chaudhary, Praful Hambarde 외 arxiv

Large language models (LLMs) achieve impressive performance when a task is fully specified in a single turn, yet the same models lose up to 39% of that performance when the identical task is revealed incrementally across…

Reinforcement Learning

Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL

2026-06-11 · Shu Tong Luo, Wenqin Liu, Rui Liu, Mingming Gong 외 arxiv

When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially …

A Large-Scale Chinese Short-Text Conversation Dataset

2020-08-10 · Yida Wang, Pei Ke, Yinhe Zheng, Kaili Huang 외

The advancements of neural dialogue generation models show promising results on modeling short-text conversations. However, training such models usually needs a large-scale high-quality dialogue corpus, which is hard to …

Dialogue GenerationShort-Text Conversation

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards

2025-10-21 · Ming Li, Pei Chen, Zhenhao Zhang, Tao Yang 외 arxiv

Large Language Models demonstrate strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC), a degradation in performance as information is revealed progressively in multi-turn s…

Reinforcement LearningInstruction Following

MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation

2026-04-09 · Jyotika Singh, Fang Tu, Miguel Ballesteros, Weiyi Sun 외 arxiv

Large language models (LLMs) suffer significant performance degradation when user instructions and context are distributed over multiple conversational turns, yet multi-turn (MT) interactions dominate chat interfaces. Th…