paper-with-me

Papers

Iterative Foundation Model Fine-Tuning on Multiple Rewards

2025-10-31 · Pouya M. Ghari, Simone Sciabola, Ye Wang arxiv

Fine-tuning foundation models has emerged as a powerful approach for generating objects with specific desired properties. Reinforcement learning (RL) provides an effective framework for this purpose, enabling models to generate outputs that maximize a given reward function. However, in many applications such as text generation and drug discovery, it can be suboptimal to optimize using a single reward signal, as multiple evaluation criteria are often necessary. This paper proposes a novel reinforcement learning-based method for fine-tuning foundation models using multiple reward signals. By employing an iterative fine-tuning strategy across these rewards, our approach generalizes state-of-the-art RL-based methods. We further provide a theoretical analysis that offers insights into the performance of multi-reward RL fine-tuning. Experimental results across diverse domains including text, biological sequence, and small molecule generation, demonstrate the effectiveness of the proposed algorithm compared to state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2511.00220

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningText GenerationDrug Discovery

Similar Papers 제목 키워드 기반

Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment

2024-02-15 · Rui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu 외

We consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is generally costly and unstable to fine-tun…

GPUReinforcement Learning (RL)

ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy

2024-03-21 · Zonghan Yang, Peng Li, Ming Yan, Ji Zhang 외

Language agents have demonstrated autonomous decision-making abilities by reasoning with foundation models. Recently, efforts have been made to train language agents for performance improvement, with multi-step reasoning…

Policy Gradient Methods

IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

2026-01-06 · Yankai Jiang, Qiaoru Li, Binlu Xu, Haoran Sun 외 arxiv

Recent research on medical MLLMs has gradually shifted its focus from image-level understanding to fine-grained, pixel-level comprehension. Although segmentation serves as the foundation for pixel-level understanding, ex…

Reinforcement LearningVisual Reasoning

Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement

2024-06-17 · Weimin Xiong, YiFan Song, Xiutian Zhao, Wenhao Wu 외

Large language model agents have exhibited exceptional performance across a range of complex interactive tasks. Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they prima…

Language ModelingLanguage ModellingLarge Language Model

WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving

2025-12-22 · Pengxuan Yang, Ben Lu, Zhongpu Xia, Chao Han 외 arxiv

Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented represe…

Self-Supervised LearningRepresentation LearningReinforcement LearningAutonomous Driving