paper-with-me

홈 › Papers

Unlocking Aha Moments via Reinforcement Learning: Advancing Collaborative Visual Comprehension and Generation

2025-06-02 · Kaihang Pan, Yang Wu, Wendong Bu, Kai Shen, Juncheng Li, Yingting Wang, Yunfei Li, Siliang Tang, Jun Xiao, Fei Wu, Hang Zhao, Yueting Zhuang

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same model. Consequently, visual comprehension does not enhance visual generation, and the reasoning mechanisms of LLMs have not been fully integrated to revolutionize image generation. In this paper, we propose to enable the collaborative co-evolution of visual comprehension and generation, advancing image generation into an iterative introspective process. We introduce a two-stage training approach: supervised fine-tuning teaches the MLLM with the foundational ability to generate genuine CoT for visual generation, while reinforcement learning activates its full potential via an exploration-exploitation trade-off. Ultimately, we unlock the Aha moment in visual generation, advancing MLLMs from text-to-image tasks to unified image generation. Extensive experiments demonstrate that our model not only excels in text-to-image generation and image editing, but also functions as a superior image semantic evaluator with enhanced visual comprehension capabilities. Project Page: https://janus-pro-r1.github.io.

📄 PDF Abstract BibTeX arXiv:2506.01480

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Identification and Classification of the Most Important Moments in Students' Collaborative Chats

2017-09-01 · RANLP 2017 9 · Costin Chiru, Remus Decea

In this paper, we present an application for the automatic identification of the important moments that might occur during students{'} collaborative chats. The moments are detected based on the input received from the us…

General Classification

RayFusion: Ray Fusion Enhanced Collaborative Visual Perception

2025-10-09 · Shaohong Wang, Bin Lu, Xinyu Xiao, Hanzhi Zhong 외 arxiv

Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit de…

3D Object DetectionAutonomous DrivingDepth Estimation

Collaborative Text-to-Image Generation via Multi-Agent Reinforcement Learning and Semantic Fusion

2025-10-12 · Jiabao Shi, Minfeng Qi, Lefeng Zhang, Di Wang 외 arxiv

Multimodal text-to-image generation remains constrained by the difficulty of maintaining semantic alignment and professional-level detail across diverse visual domains. We propose a multi-agent reinforcement learning fra…

Multi-agent Reinforcement LearningText-to-Image GenerationContrastive LearningSemantic Similarity

Unlocking Feature Visualization for Deep Network with MAgnitude Constrained Optimization

2023-09-21 · NeurIPS 2023 11

Feature visualization has gained significant popularity as an explainability method, particularly after the influential work by Olah et al. in 2017. Despite its success, its widespread adoption has been limited due to is…

UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets

2025-09-18 · Pengyu Wang, Shaojun Zhou, Chenkun Tan, Xinghao Wang 외 arxiv

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image s…

Visual Question Answering