paper-with-me

홈 › Papers

ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation

2026-01-20 · Zhebo Wang, Xiaohu Mu, Zijie Zhou, Mohan Li, Wenpeng Xing, Dezhang Kong, Meng Han arxiv

Large Language Models (LLMs) in multi-turn conversations often suffer from a ``lost-in-conversation'' phenomenon, where they struggle to recover from early incorrect assumptions, particularly when users provide ambiguous initial instructions. We find that standard post-training techniques like Reinforcement Learning with Verifiable Rewards (RLVR) exacerbate this issue by rewarding confident, direct answers, thereby inducing overconfidence and discouraging the model from seeking clarification. To address this, we propose Illocution-Calibrated Policy Optimization (ICPO), a novel training framework that sensitizes the model to instruction ambiguity. ICPO augments the training corpus with underspecified prompts and conditions the reward signal on the user's illocutionary intent, rewarding the model for expressing uncertainty or asking for clarification when faced with ambiguity. Experiments demonstrate that ICPO fosters appropriate humility, yielding a substantial average improvement of 75\% in multi-turn conversation, while preserving robust performance on single-turn benchmarks. Our work presents a practical path toward more robust and collaborative conversational AI that can better navigate the nuances of human interaction.

📄 PDF Abstract BibTeX arXiv:2601.15330

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Provable and Practical In-Context Policy Optimization for Self-Improvement

2026-03-02 · Tianrun Yu, Yuxiao Yang, Zhaoyang Wang, Kaixiang Zhao 외 arxiv

We study test-time scaling, where a model improves its answer through multi-round self-reflection at inference. We introduce In-Context Policy Optimization (ICPO), in which an agent optimizes its response in context usin…

Mathematical Reasoning

Think Outside the Policy: In-Context Steered Policy Optimization

2025-10-30 · Hsiu-Yuan Huang, Chenming Tang, Weijie Liu, Clive Bai 외 arxiv

Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable progress in improving the reasoning capabilities of Large Reasoning Mode…

Reinforcement LearningMathematical Reasoning

SocraticPO: Policy Optimization via Interactive Guidance

2026-06-03 · Zirui Liu, Jie Ouyang, Qi Liu, Xianquan Wang 외 arxiv

Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization direction but rarely explain how a model sh…

Reinforcement Learning

DynamicPO: Dynamic Preference Optimization for Recommendation

2026-05-01 · Xingyu Hu, Kai Zhang, Jiancan Wu, Shuli Wang 외 arxiv

In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with user preferences, requiring multi-negative objective functions to leverage abundant…

Recommendation Systems

ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning

2025-11-26 · Jinpeng Wang, Chao Li, Ting Ye, Mengyuan Zhang 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing RLVR methods are often constrained by is…

Reinforcement Learning