paper-with-me

홈 › Papers

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

2025-05-22 · Ilgee Hong, Changlong Yu, Liang Qiu, Weixiang Yan, Zhenghao Xu, Haoming Jiang, Qingru Zhang, Qin Lu, Xin Liu, Chao Zhang, Tuo Zhao

Reinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from sensitivity to data size and coverage, as well as vulnerability to reward hacking. Generative reward models (GenRMs) offer a more robust alternative by generating chain-of-thought (CoT) rationales followed by a final reward. However, existing GenRMs rely on shallow, vertically scaled reasoning, limiting their capacity to handle nuanced or complex (e.g., reasoning-intensive) tasks. Moreover, their pairwise preference outputs are incompatible with standard RLHF algorithms that require pointwise reward signals. In this work, we introduce Think-RM, a training framework that enables long-horizon reasoning in GenRMs by modeling an internal thinking process. Rather than producing structured, externally provided rationales, Think-RM generates flexible, self-guided reasoning traces that support advanced capabilities such as self-reflection, hypothetical reasoning, and divergent reasoning. To elicit these reasoning abilities, we first warm-up the models by supervised fine-tuning (SFT) over long CoT data. We then further improve the model's long-horizon abilities by rule-based reinforcement learning (RL). In addition, we propose a novel pairwise RLHF pipeline that directly optimizes policies using pairwise preference rewards, eliminating the need for pointwise reward conversion and enabling more effective use of Think-RM outputs. Experiments show that Think-RM achieves state-of-the-art results on RM-Bench, outperforming both BT RM and vertically scaled GenRM by 8%. When combined with our pairwise RLHF pipeline, it demonstrates superior end-policy performance compared to traditional approaches.

📄 PDF Abstract BibTeX arXiv:2505.16265

Code (1)

ilgeehong/think-rm 공식 구현 pytorch

Tasks

Reinforcement Learning (RL)

Similar Papers 제목 키워드 기반

DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models

2025-12-30 · Zefeng He, Xiaoye Qu, Yafu Li, Tong Zhu 외 arxiv

While recent Multimodal Large Language Models (MLLMs) have attained significant strides in multimodal reasoning, their reasoning processes remain predominantly text-centric, leading to suboptimal performance in complex l…

Multimodal Reasoning

Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization

2026-02-26 · Qianben Chen, Tianrui Qin, King Zhu, Qiexiang Wang 외 arxiv

Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios. Moreover, generalization across heterogeneous researc…

Reinforcement LearningQuestion Answering

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

2026-04-10 · Lulin Liu, Dayou Li, Yiqing Liang, Sicong Jiang 외 arxiv

Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household tasks. However, VLM-based auto-labeling is often noisy because the pri…

Instruction Following

ThinkingVLA: Interleaved Vision and Language Reasoning for Robotic Manipulation

2026-06-16 · Tianyi Lu, Hui Zhang, Zijie Diao, Junke Wang 외 arxiv

Most Vision-Language-Action (VLA) models map observations directly to actions without explicit reasoning, limiting their capacity for reasoning-intensive long-horizon tasks. To address this, existing approaches adopt Cha…

Visual Reasoning

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

2026-02-06 · Yuchen Yan, Liang Jiang, Jin Jiang, Shuaicheng Li 외 arxiv

Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, context length limits, and degraded reasoning due to lost-in-the-middle effects…

Reinforcement Learning