paper-with-me

홈 › Papers

Enhancing LLM Reasoning with Iterative DPO: A Comprehensive Empirical Investigation

2025-03-17 · Songjun Tu, Jiahao Lin, Xiangyu Tian, Qichao Zhang, Linjing Li, Yuqian Fu, Nan Xu, wei he, Xiangyuan Lan, Dongmei Jiang, Dongbin Zhao

Recent advancements in post-training methodologies for large language models (LLMs) have highlighted reinforcement learning (RL) as a critical component for enhancing reasoning. However, the substantial computational costs associated with RL-based approaches have led to growing interest in alternative paradigms, such as Direct Preference Optimization (DPO). In this study, we investigate the effectiveness of DPO in facilitating self-improvement for LLMs through iterative preference-based learning. We demonstrate that a single round of DPO with coarse filtering significantly enhances mathematical reasoning performance, particularly for strong base model. Furthermore, we design an iterative enhancement framework for both the generator and the reward model (RM), enabling their mutual improvement through online interaction across multiple rounds of DPO. Finally, with simple verifiable rewards, our model DPO-VP achieves RL-level performance with significantly lower computational overhead. These findings highlight DPO as a scalable and cost-effective alternative to RL, offering a practical solution for enhancing LLM reasoning in resource-constrained situations.

📄 PDF Abstract BibTeX arXiv:2503.12854

Code (1)

TU2021/DPO-VP 공식 구현 pytorch

Tasks

Mathematical ReasoningReinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

BASE 설명 없음
DPO 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

2025-10-02 · Qiyuan Liu, Hao Xu, Xuhong Chen, Wei Chen 외 arxiv

Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune LLMs during reinforcement learning (RL) and help select the best answer …

Reinforcement Learning

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

2025-05-21 · Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang 외

Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enhanced reasoning capabilities do not necessarily translate to improved saf…

Math

Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs

2025-09-28 · Shulin Huang, Yiran Ding, Junshu Pan, Yue Zhang arxiv

Enhancing the complex reasoning capabilities of Large Language Models (LLMs) attracts widespread attention. While reinforcement learning (RL) has shown superior performance for improving complex reasoning, its impact on …

Reinforcement Learning

Enhancing Chain-of-Thoughts Prompting with Iterative Bootstrapping in Large Language Models

2023-04-23 · Jiashuo Sun, Yi Luo, Yeyun Gong, Chen Lin 외

Large language models (LLMs) can achieve highly effective performance on various reasoning tasks by incorporating step-by-step chain-of-thought (CoT) prompting as demonstrations. However, the reasoning chains of demonstr…

Double-Checker: Enhancing Reasoning of Slow-Thinking LLMs via Self-Critical Fine-Tuning

2025-06-26 · Xin Xu, Tianhao Chen, Fan Zhang, Wanlong Liu 외

While slow-thinking large language models (LLMs) exhibit reflection-like reasoning, commonly referred to as the "aha moment:, their ability to generate informative critiques and refine prior solutions remains limited. In…