paper-with-me

홈 › Papers

Visual Grounding for Object-Level Generalization in Reinforcement Learning

2024-08-04 · Haobin Jiang, Zongqing Lu

Generalization is a pivotal challenge for agents following natural language instructions. To approach this goal, we leverage a vision-language model (VLM) for visual grounding and transfer its vision-language knowledge into reinforcement learning (RL) for object-centric tasks, which makes the agent capable of zero-shot generalization to unseen objects and instructions. By visual grounding, we obtain an object-grounded confidence map for the target object indicated in the instruction. Based on this map, we introduce two routes to transfer VLM knowledge into RL. Firstly, we propose an object-grounded intrinsic reward function derived from the confidence map to more effectively guide the agent towards the target object. Secondly, the confidence map offers a more unified, accessible task representation for the agent's policy, compared to language embeddings. This enables the agent to process unseen objects and instructions through comprehensible visual confidence maps, facilitating zero-shot object-level generalization. Single-task experiments prove that our intrinsic reward significantly improves performance on challenging skill learning. In multi-task experiments, through testing on tasks beyond the training set, we show that the agent, when provided with the confidence map as the task representation, possesses better generalization capabilities than language-based conditioning. The code is available at https://github.com/PKU-RL/COPL.

📄 PDF Abstract BibTeX arXiv:2408.01942

Code (1)

pku-rl/copl 공식 구현 pytorch

Tasks

Language ModellingObjectreinforcement-learningReinforcement LearningReinforcement Learning (RL)Visual GroundingZero-shot Generalization

Similar Papers 제목 키워드 기반

STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning

2026-02-12 · Xiaowen Zhang, Zhi Gao, Licheng Jiao, Lingling Li 외 arxiv

In vision-language models (VLMs), misalignment between textual descriptions and visual coordinates often induces hallucinations. This issue becomes particularly severe in dense prediction tasks such as spatial-temporal v…

Referring Video Object SegmentationZero-shot GeneralizationReinforcement LearningVideo Grounding

MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision

2025-08-11 · Zhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing 외 arxiv

Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current…

Reinforcement Learning

MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning

2025-09-26 · Lihao Zheng, Jiawei Chen, Xintian Shen, Hao Ma 외 arxiv

Multi-image reasoning and grounding require understanding complex cross-image relationships at both object levels and image levels. Current Large Visual Language Models (LVLMs) face two critical challenges: the lack of c…

Reinforcement Learning

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

2025-06-05 · Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan 외

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplored. In this paper, we introduce Anywhere3…

3D visual groundingObjectReferring ExpressionSpatial Reasoning+1

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

2026-07-16 · Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Chia-Wei Tang 외 arxiv

Multimodal large language models (MLLMs) ground whole objects well from free-form language queries, but they struggle when the query names a part rather than the object. We trace this to a missing object-part hierarchy, …

Reinforcement LearningVisual Grounding