paper-with-me

홈 › Papers

AdaGRPO: A Capability-Aware Adaptive Enhancement for Flow-based GRPO

2026-06-05 · Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Tianyi Wei, Xiaohang Zhan, Jiaqi Wang, Tong Wu, Xingang Pan, Dahua Lin arxiv

Group Relative Policy Optimization (GRPO) has demonstrated remarkable success in aligning text-to-image (T2I) flow models with human preferences. However, we have identified that the learning loop of current flow-based GRPO is fundamentally decoupled from the learner's current capability, suffering from critical blind spots at both prompt selection and advantage estimation: (i) Existing methods sample prompts randomly, overlooking the substantial impact of data selection on reinforcement learning (RL) efficacy--a factor proven crucial in GRPO for large language models; (ii) They evaluate sample quality solely relying on intra-group statistics, lacking a global perspective to accurately measure true policy improvement. To address these issues, we propose Adaptive GRPO (AdaGRPO), a novel capability-aware RL algorithm tailored for flow models. Specifically, AdaGRPO consists of two principal components: (i) Online Curriculum Filtering Strategy: Dynamically tracks the model's proficiency and adaptively selects prompts that best match its current learning boundary; (ii) Cross-Level Advantage Fusion: Synergistically integrates fine-grained intra-group advantages with macro-level global advantages, providing a comprehensive and unbiased policy evaluation. As a lightweight, plug-and-play module, AdaGRPO can be seamlessly integrated with existing frameworks such as Flow-GRPO, DanceGRPO, and Flow-CPS. Extensive experiments demonstrate that AdaGRPO consistently drives performance gains while significantly stabilizes GRPO training for flow models.

📄 PDF Abstract BibTeX arXiv:2606.06828

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning

2025-09-26 · Zejun Li, Yingxiu Zhao, Jiwen Zhang, Siyuan Wang 외 arxiv

Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be achieved in particular domains, they struggle to develop general reasoning capabilities. Inspired by this,…

Visual Reasoning

MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents

2025-09-10 · Yifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu 외 arxiv

Building general-purpose graphical user interface (GUI) agents has become increasingly promising with the progress in vision language models. However, developing effective mobile GUI agents with reinforcement learning (R…

Reinforcement Learning

Adaptive Loss Balancing for Noise-Robust GRPO in Generative Recommendation

2026-06-07 · Kewei Xu, Junbo Qi, Yanyan Zou, Pengfei Zhang 외 arxiv

Reinforcement learning (RL) presents a promising avenue for enhancing generative recommendation beyond supervised imitation, leveraging reward signals to guide policy improvement. However, its efficacy is critically cont…

Reinforcement Learning

Zero-TIG: Temporal Consistency-Aware Zero-Shot Illumination-Guided Low-light Video Enhancement

2025-03-14 · Yini Li, Nantheera Anantrasirichai

Low-light and underwater videos suffer from poor visibility, low contrast, and high noise, necessitating enhancements in visual quality. However, existing approaches typically rely on paired ground truth, which limits th…

DenoisingImage DenoisingOptical Flow EstimationVideo Enhancement+1

FlowLUT: Efficient Image Enhancement via Differentiable LUTs and Iterative Flow Matching

2025-09-28 · Liubing Hu, Chen Wu, Anrui Wang, Dianjie Lu 외 arxiv

Deep learning-based image enhancement methods face a fundamental trade-off between computational efficiency and representational capacity. For example, although a conventional three-dimensional Look-Up Table (3D LUT) can…

Computational EfficiencyImage Enhancement