paper-with-me

홈 › Papers

AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning

2025-08-09 · Shihao Yuan, Yahui Liu, Yang Yue, Jingyuan Zhang, Wangmeng Zuo, Qi Wang, Fuzheng Zhang, Guorui Zhou arxiv

Inspired by the success of reinforcement learning (RL) in refining large language models (LLMs), we propose AR-GRPO, an approach to integrate online RL training into autoregressive (AR) image generation models. We adapt the Group Relative Policy Optimization (GRPO) algorithm to refine the vanilla autoregressive models' outputs by carefully designed reward functions that evaluate generated images across multiple quality dimensions, including perceptual quality, realism, and semantic fidelity. We conduct comprehensive experiments on both class-conditional (i.e., class-to-image) and text-conditional (i.e., text-to-image) image generation tasks, demonstrating that our RL-enhanced framework significantly improves both the image quality and human preference of generated images compared to the standard AR baselines. Our results show consistent improvements across various evaluation metrics, establishing the viability of RL-based optimization for AR image generation and opening new avenues for controllable and high-quality image synthesis. The source codes and models are available at: https://github.com/Kwai-Klear/AR-GRPO.

📄 PDF Abstract BibTeX arXiv:2508.06924

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningImage Generation

Similar Papers 제목 키워드 기반

Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training

2026-06-19 · Yuanhao Chiang, Hongbo Duan, Chunru Yang, Jiahua Pei 외 arxiv

Autoregressive text-to-image (T2I) generation has recently advanced rapidly, yet aligning generated images with human preferences remains challenging. GRPO-style online reinforcement learning provides an effective framew…

Reinforcement Learning

STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation

2025-09-29 · Xiaoxiao Ma, Haibo Qiu, Guohui Zhang, Zhixiong Zeng 외 arxiv

Reinforcement learning has recently been explored to improve text-to-image generation, yet applying existing GRPO algorithms to autoregressive (AR) image models remains challenging. The instability of the training proces…

Text-to-Image GenerationReinforcement Learning

UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation

2026-03-24 · Jie Liu, Zilyu Ye, Linxiao Yuan, Shenhan Zhu 외 arxiv

Unified models capable of interleaved generation have emerged as a promising paradigm, with the community increasingly converging on autoregressive modeling for text and flow matching for image generation. To advance thi…

Reinforcement Learningmultimodal generationImage Generation

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

2025-10-15 · Yifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun 외 arxiv

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an impor…

Text-to-Image GenerationReinforcement Learning

Advances in GRPO for Generation Models: A Survey

2026-02-21 · Zexiang Liu, Xianglong He, Yangguang Li arxiv

Large-scale flow matching models have achieved strong performance across generative tasks such as text-to-image, video, 3D, and speech synthesis. However, aligning their outputs with human preferences and task-specific o…

Reinforcement LearningSpeech SynthesisVideo GenerationImage Editing