paper-with-me

홈 › Papers

Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL

2026-06-11 · Shu Tong Luo, Wenqin Liu, Rui Liu, Mingming Gong, Jiaxian Guo arxiv

When a user reveals task-critical information across several conversation turns, LLM accuracy drops by up to 65% despite full context availability. We show that this Lost in Conversation degradation can be substantially mitigated by training models to maintain a compact rolling memory instead of attending to a growing history. To make such training scalable, we introduce a low-cost sharding pipeline that converts single-turn QA datasets into multi-turn fragmented-information episodes, eliminating the need for hours of manual annotation. Training only on sharded GSM8K, our memory-augmented policy significantly improves multi-turn accuracy and generalises zero-shot to harder math and out-of-domain long-context QA. Moreover, memory-trained models outperform full-history baselines even when given the full history at test time, suggesting that learning to compress induces more robust incremental reasoning than full-context exposure alone.

📄 PDF Abstract BibTeX arXiv:2606.12941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

2026-08-13 · Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi 외 hf

LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is confined to the Thought phase: while the agent serializes an action and waits for the environment, its reas…

Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers

2026-05-13 · Victor Norgren arxiv

Conventional transformer inference engines are request-driven, paying an O(n) prefill cost on every query. In streaming workloads, where data arrives continuously and queries probe an ever-growing context, this cost is p…

A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning

2025-07-18 · Licheng Liu, Zihan Wang, Linjie Li, Chenwei Xu 외 arxiv

Multi-turn problem solving is critical yet challenging for Large Reasoning Models (LRMs) to reflect on their reasoning and revise from feedback. Existing Reinforcement Learning (RL) methods train large reasoning models o…

Reinforcement Learning

Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction

2026-03-05 · Xingwu Chen, Zhanqiu Zhang, Yiwen Guo, Difan Zou arxiv

While LLMs demonstrate strong reasoning capabilities when provided with full information in a single turn, they exhibit substantial vulnerability in multi-turn interactions. Specifically, when information is revealed inc…

Reinforcement LearningDomain Generalization

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

2026-08-29 · Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan 외 hf

Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents…