paper-with-me

홈 › Papers

A Study on Dialogue Reward Prediction for Open-Ended Conversational Agents

2018-12-02 · Heriberto Cuayáhuitl, Seonghan Ryu, Donghyeon Lee, Jihie Kim

The amount of dialogue history to include in a conversational agent is often underestimated and/or set in an empirical and thus possibly naive way. This suggests that principled investigations into optimal context windows are urgently needed given that the amount of dialogue history and corresponding representations can play an important role in the overall performance of a conversational system. This paper studies the amount of history required by conversational agents for reliably predicting dialogue rewards. The task of dialogue reward prediction is chosen for investigating the effects of varying amounts of dialogue history and their impact on system performance. Experimental results using a dataset of 18K human-human dialogues report that lengthy dialogue histories of at least 10 sentences are preferred (25 sentences being the best in our experiments) over short ones, and that lengthy histories are useful for training dialogue reward predictors with strong positive correlations between target dialogue rewards and predicted ones.

📄 PDF Abstract BibTeX arXiv:1812.00350

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

2025-10-17 · Pengkai Wang, Pengwei Liu, Qi Zuo, Zhijie Sang 외 arxiv

Reinforcement learning (RL) has powered many recent breakthroughs in large language models (LLMs), especially for tasks where rewards can be computed automatically, such as code generation. However, it is less effective …

Reinforcement LearningCode Generation

Tailored Conversations beyond LLMs: A RL-Based Dialogue Manager

2025-06-24 · Lucie Galland, Catherine Pelachaud, Florian Pecune

In this work, we propose a novel framework that integrates large language models (LLMs) with an RL-based dialogue manager for open-ended dialogue with a specific goal. By leveraging hierarchical reinforcement learning to…

Hierarchical Reinforcement LearningMeta-Learning

RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems

2025-12-11 · Hang Ding, Qiming Feng, Dongqi Liu, Qi Zhao 외 arxiv

Reward modeling has become a cornerstone of aligning large language models (LLMs) with human preferences. Yet, when extended to subjective and open-ended domains such as role play, existing reward models exhibit severe d…

Dialogue Model Optimization via Agent Game and Adaptive Tree-based GRPO

2026-02-09 · Kun Peng, Conghui Tan, Yu Liu, Guohua Tang 외 arxiv

Open-ended dialogue agents aim to deliver engaging, personalized interactions by adapting to users' traits, but existing methods face critical limitations: over-reliance on pre-collected user data, and short-horizon bias…

Reinforcement Learning

CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization

2025-08-12 · Xinge Ye, Rui Wang, Yuchuan Wu, Victor Ma 외 arxiv

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles with open-ended subjective tasks like rol…

Reinforcement LearningMathematical ReasoningCode Generation