paper-with-me

홈 › Papers

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

2026-05-15 · Yuyuan Liu, Yiping Ji, Anjie Le, Jiayuan Zhu, Jiazhen Pan, Can Peng, Jiajun Deng, Fengbei Liu, Junde Wu arxiv

Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing methods, mainly based on GRPO, assign rewards at the response level. Such sparse reward, often criterion-induced, leads to minimal learning signals when all candidate responses fail in challenging scenarios. In this work, we propose a group-revision optimisation paradigm that enhances learning on hard cases. It begins with a sampled initial response and generates a set of revised candidates to explore improved grounding outcomes. Inspired by reward shaping, we introduce a consolidation process that quantifies each candidate's improvement over the initial attempt and converts it into informative shaping signals. These signals are used to both refine the reward and modulate the advantage, amplifying the influence of high-quality revisions. Our method achieves consistent gains across referring and reasoning segmentation, REC, and counting benchmarks compared with prior GRPO-based models. Our code is available at https://github.com/yyliu01/GroupRevision.

📄 PDF Abstract BibTeX arXiv:2605.15951

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning

2026-01-28 · Minwu Kim, Safal Shrestha, Anubhav Shrestha, Keith Ross arxiv

As Reinforcement Learning with Verifiable Rewards (RLVR) substantially improves the reasoning abilities of large language models (LLMs), a new bottleneck emerges: more training problems become saturated, that is, the LLM…

Reinforcement Learning

Unleashing VLA Potentials in Autonomous Driving via Explicit Learning from Failures

2026-03-01 · Yuechen Luo, Qimao Chen, Fang Li, Shaoqing Xu 외 arxiv

Vision-Language-Action (VLA) models for autonomous driving often hit a performance plateau during Reinforcement Learning (RL) optimization. This stagnation arises from exploration capabilities constrained by previous Sup…

Reinforcement LearningAutonomous Driving

Understanding Teacher Revisions of Large Language Model-Generated Feedback

2026-03-29 · Conrad Borchers, Luiz Rodrigues, Newarney Torrezão da Costa, Cleon Xavier 외 arxiv

Large language models (LLMs) increasingly generate formative feedback for students, yet little is known about how teachers revise this feedback before it reaches learners. Teachers' revisions shape what students receive,…

eRevise+RF: A Writing Evaluation System for Assessing Student Essay Revisions and Providing Formative Feedback

2025-01-01 · Zhexiong Liu, Diane Litman, Elaine Wang, Tianwen Li 외

The ability to revise essays in response to feedback is important for students' writing success. An automated writing evaluation (AWE) system that supports students in revising their essays is thus essential. We present …

Automated Writing Evaluation

Integrating AI for Enhanced Feedback in Translation Revision- A Mixed-Methods Investigation of Student Engagement

2024-10-11 · Simin Xu, Yanfang Su, Kanglong Liu

Despite the well-established importance of feedback in education, the application of Artificial Intelligence (AI)-generated feedback, particularly from language models like ChatGPT, remains understudied in translation ed…

Translation