paper-with-me

홈 › Papers

OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models

2026-04-05 · Liyu Zhang, Kehan Li, Tingrui Han, Tao Zhao, Yuxuan Sheng, Shibo He, Chao Li arxiv

Post training via GRPO has demonstrated remarkable effectiveness in improving the generation quality of flow-matching models. However, GRPO suffers from inherently low sample efficiency due to its on-policy training paradigm. To address this limitation, we present OP-GRPO, the first Off-Policy GRPO framework tailored for flow-matching models. First, we actively select high-quality trajectories and adaptively incorporate them into a replay buffer for reuse in subsequent training iterations. Second, to mitigate the distribution shift introduced by off-policy samples, we propose a sequence-level importance sampling correction that preserves the integrity of GRPO's clipping mechanism while ensuring stable policy updates. Third, we theoretically and empirically show that late denoising steps yield ill-conditioned off-policy ratios, and mitigate this by truncating trajectories at late steps. Across image and video generation benchmarks, OP-GRPO achieves comparable or superior performance to Flow-GRPO with only 34.2% of the training steps on average, yielding substantial gains in training efficiency while maintaining generation quality.

📄 PDF Abstract BibTeX arXiv:2604.04142

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

2025-10-24 · Yifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du 외 arxiv

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccu…

Reinforcement Learning

Reinforcement Learning for Flow-Matching Policies

2025-07-20 · Samuel Pfrommer, Yixiao Huang, Somayeh Sojoudi arxiv

Flow-matching policies have emerged as a powerful paradigm for generalist robotics. These models are trained to imitate an action chunk, conditioned on sensor observations and textual instructions. Often, training demons…

Reinforcement Learning

Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models

2025-11-21 · Dailan He, Guanlin Feng, Xingtong Ge, Yazhe Niu 외 arxiv

Group Relative Policy Optimization (GRPO) has shown promise in aligning image and video generative models with human preferences. However, applying it to modern flow matching models is challenging because of its determin…

Computational EfficiencyContrastive Learning

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

2026-06-05 · Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang 외 arxiv

Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that the learning loop of current flow-based G…

Reinforcement Learning

Advances in GRPO for Generation Models: A Survey

2026-02-21 · Zexiang Liu, Xianglong He, Yangguang Li arxiv

Large-scale flow matching models have achieved strong performance across generative tasks such as text-to-image, video, 3D, and speech synthesis. However, aligning their outputs with human preferences and task-specific o…

Reinforcement LearningSpeech SynthesisVideo GenerationImage Editing