paper-with-me

홈 › Papers

Plug-and-Play Training Framework for Preference Optimization

2024-12-30 · Jingyuan Ma, Rui Li, Zheng Li, Lei Sha, Zhifang Sui

Recently, preference optimization methods such as DPO have significantly enhanced large language models (LLMs) in wide tasks including dialogue and question-answering. However, current methods fail to account for the varying difficulty levels of training samples during preference optimization, leading to mediocre performance in tasks with high accuracy requirements, particularly in mathematical reasoning. To address this limitation, we propose a novel training framework, which employs multiple sampling to analyze output distributions, assign different weights to samples, and incorporate these weights into the preference optimization process. This plug-and-play approach enables LLMs to prioritize challenging examples during training, improving learning efficiency. Experimental results demonstrate that our framework integrates seamlessly with various preference optimization methods and achieves consistent improvements in mathematical reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2412.20996

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical ReasoningQuestion Answering

Methods 이 논문이 사용한 방법론

DPO 설명 없음

Similar Papers 제목 키워드 기반

Toward Preference-aligned Large Language Models via Residual-based Model Steering

2025-09-28 · Lucio La Cava, Andrea Tagarelli arxiv

Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Opti…

Reinforcement LearningMathematical ReasoningCode Generation

Sample-efficient LLM Optimization with Reset Replay

2025-08-08 · Zichuan Liu, Jinyu Wang, Lei Song, Jiang Bian arxiv

Recent advancements in LLM post-training, particularly through reinforcement learning and preference optimization, are key to boosting their reasoning capabilities. However, these methods often suffer from low sample eff…

Reinforcement Learning

UCPO: A Universal Constrained Combinatorial Optimization Method via Preference Optimization

2025-11-13 · Zhanhong Fang, Debing Wang, Jinbiao Chen, Jiahai Wang 외 arxiv

Neural solvers have demonstrated remarkable success in combinatorial optimization, often surpassing traditional heuristics in speed, solution quality, and generalization. However, their efficacy deteriorates significantl…

X-ALMA: Plug & Play Modules and Adaptive Rejection for Quality Translation at Scale

2024-10-04 · Haoran Xu, Kenton Murray, Philipp Koehn, Hieu Hoang 외

Large language models (LLMs) have achieved remarkable success across various NLP tasks, yet their focus has predominantly been on English due to English-centric pre-training and limited multilingual data. While some mult…

Machine TranslationTranslation

DynamicPO: Dynamic Preference Optimization for Recommendation

2026-05-01 · Xingyu Hu, Kai Zhang, Jiancan Wu, Shuli Wang 외 arxiv

In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with user preferences, requiring multi-negative objective functions to leverage abundant…

Recommendation Systems