paper-with-me

홈 › Papers

Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design

2026-02-25 · Haoxiang Sun, Tao Wang, Chenwei Tang, Li Yuan, Jiancheng Lv arxiv

Following the success of Group Relative Policy Optimization (GRPO) in foundation LLMs, an increasing number of works have sought to adapt GRPO to Visual Large Language Models (VLLMs) for visual perception tasks (e.g., detection and segmentation). However, much of this line of research rests on a long-standing yet unexamined assumption: training paradigms developed for language reasoning can be transferred seamlessly to visual perception. Our experiments show that this assumption is not valid, revealing intrinsic differences between reasoning-oriented and perception-oriented settings. Using reasoning segmentation as a representative case, we surface two overlooked factors: (i) the need for a broader output space, and (ii) the importance of fine-grained, stable rewards. Building on these observations, we propose Dr.~Seg, a simple, plug-and-play GRPO-based framework consisting of a Look-to-Confirm mechanism and a Distribution-Ranked Reward module, requiring no architectural modifications and integrating seamlessly with existing GRPO-based VLLMs. Extensive experiments demonstrate that Dr.~Seg improves performance in complex visual scenarios while maintaining strong generalization. Code, models, and datasets are available at https://github.com/eVI-group-SCU/Dr-Seg.

📄 PDF Abstract BibTeX arXiv:2603.00152

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

2025-05-28 · Youssef Mroueh, Nicolas Dupuis, Brian Belgodere, Apoorva Nitsure 외

We revisit Group Relative Policy Optimization (GRPO) in both on-policy and off-policy optimization regimes. Our motivation comes from recent work on off-policy Proximal Policy Optimization (PPO), which improves training …

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning

2026-05-05 · Yujun Li, Hongyuan Zhang, Yuan Yuan arxiv

Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models. It raises a question of whether the GRPO also significantly promotes the …

Reinforcement LearningTest-time Adaptation

Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View

2025-11-10 · Jianyu Qi, Ding Zou, Wenrui Yan, Rui Ma 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have spurred significant progress in Chain-of-Thought (CoT) reasoning. Building on the success of Deepseek-R1, researchers extended multimodal reasoning to post…

Reinforcement LearningMultimodal Reasoning

AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO

2025-02-20 · Alan Dao, Dinh Bach Vu

Large Language Models (LLMs) have demonstrated impressive capabilities in language processing, yet they often struggle with tasks requiring genuine visual spatial reasoning. In this paper, we introduce a novel two-stage …

Autonomous NavigationNavigateSequential Decision MakingSpatial Reasoning+1

Improved Visual-Spatial Reasoning via R1-Zero-Like Training

2025-04-01 · Zhenyi Liao, Qingsong Xie, Yanhao Zhang, Zijian Kong 외

Increasing attention has been placed on improving the reasoning capacities of multi-modal large language models (MLLMs). As the cornerstone for AI agents that function in the physical realm, video-based visual-spatial in…

GPUSpatial Reasoning