paper-with-me

Papers

Proximal Supervised Fine-Tuning

2025-08-25 · Wenhong Zhu, Ruobing Xie, Rui Wang, Xingwu Sun, Di Wang, Pengfei Liu arxiv

Supervised fine-tuning (SFT) of foundation models often leads to poor generalization, where prior capabilities deteriorate after tuning on new tasks or domains. Inspired by trust-region policy optimization (TRPO) and proximal policy optimization (PPO) in reinforcement learning (RL), we propose Proximal SFT (PSFT). This fine-tuning objective incorporates the benefits of trust-region, effectively constraining policy drift during SFT while maintaining competitive tuning. By viewing SFT as a special case of policy gradient methods with constant positive advantages, we derive PSFT that stabilizes optimization and leads to generalization, while leaving room for further optimization in subsequent post-training stages. Experiments across mathematical and human-value domains show that PSFT matches SFT in-domain, outperforms it in out-of-domain generalization, remains stable under prolonged training without causing entropy collapse, and provides a stronger foundation for the subsequent optimization.

📄 PDF Abstract BibTeX arXiv:2508.17784

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDomain Generalization

Similar Papers 제목 키워드 기반

PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models

2024-02-13 · CVPR 2024 1 · Fei Deng, Qifei Wang, Wei Wei, Matthias Grundmann 외

Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize…

DenoisingReinforcement Learning (RL)

Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

2024-05-28 · Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin 외

Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data. First time in the literature, we show that the jail-broken effect can be…

Safety Alignment

Is Crowdsourcing Breaking Your Bank? Cost-Effective Fine-Tuning of Pre-trained Language Models with Proximal Policy Optimization

2024-02-28 · Shuo Yang, Gjergji Kasneci

Wide usage of ChatGPT has highlighted the potential of reinforcement learning from human feedback. However, its training pipeline relies on manual ranking, a resource-intensive process. To reduce labor costs, we propose …

Language ModelingLanguage Modelling

The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

2024-08-23 · Venkatesh Balavadhani Parthasarathy, Ahtsham Zafar, Aafaq Khan, Arsalan Shahid

This report examines the fine-tuning of Large Language Models (LLMs), integrating theoretical insights with practical applications. It outlines the historical evolution of LLMs from traditional Natural Language Processin…

Computational EfficiencyInference OptimizationMixture-of-Experts

$Σ$-net: Ensembled Iterative Deep Neural Networks for Accelerated Parallel MR Image Reconstruction

2019-12-11 · Jo Schlemper, Chen Qin, Jinming Duan, Ronald M. Summers 외

We explore an ensembled $\Sigma$-net for fast parallel MR imaging, including parallel coil networks, which perform implicit coil weighting, and sensitivity networks, involving explicit sensitivity maps. The networks in $…

Image ReconstructionSensitivitySSIM