paper-with-me

Papers

Improving Generalized Visual Grounding with Instance-aware Joint Learning

2025-09-17 · Ming Dai, Wenxuan Cheng, Jiang-Jiang Liu, Lingfeng Yang, Zhenhua Feng, Wankou Yang, Jingdong Wang arxiv

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios. Specifically, GREC focuses on accurately identifying all referential objects at the coarse bounding box level, while GRES aims for achieve fine-grained pixel-level perception. However, existing approaches typically treat these tasks independently, overlooking the benefits of jointly training GREC and GRES to ensure consistent multi-granularity predictions and streamline the overall process. Moreover, current methods often treat GRES as a semantic segmentation task, neglecting the crucial role of instance-aware capabilities and the necessity of ensuring consistent predictions between instance-level boxes and masks. To address these limitations, we propose InstanceVG, a multi-task generalized visual grounding framework equipped with instance-aware capabilities, which leverages instance queries to unify the joint and consistency predictions of instance-level boxes and masks. To the best of our knowledge, InstanceVG is the first framework to simultaneously tackle both GREC and GRES while incorporating instance-aware capabilities into generalized visual grounding. To instantiate the framework, we assign each instance query a prior reference point, which also serves as an additional basis for target matching. This design facilitates consistent predictions of points, boxes, and masks for the same instance. Extensive experiments obtained on ten datasets across four tasks demonstrate that InstanceVG achieves state-of-the-art performance, significantly surpassing the existing methods in various evaluation metrics. The code and model will be publicly available at https://github.com/Dmmm1997/InstanceVG.

📄 PDF Abstract BibTeX arXiv:2509.13747

Code (0)

등록된 구현이 없습니다.

Tasks

Generalized Referring Expression ComprehensionSemantic SegmentationVisual Grounding

Similar Papers 제목 키워드 기반

AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding

2026-05-21 · Haocheng Li, Juepeng Zheng, Zenghao Yang, Kaiqi Du 외 arxiv

Visual grounding, the task of localizing objects described by natural-language expressions, is a foundational capability for agricultural AI systems, enabling applications such as selective weeding, disease monitoring, a…

Referring ExpressionVisual Grounding

Semantic-Spatial Discriminability Enhancement for Generalized Visual Grounding

2026-08-31 · Kaiyan Lei, Xu-Yao Zhang arxiv

Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding paradigm by integrating multi-target and non-target scenarios. Previous …

Visual Grounding

Finding "It": Weakly-Supervised Reference-Aware Visual Grounding in Instructional Videos

2018-06-01 · CVPR 2018 6 · De-An Huang, Shyamal Buch, Lucio Dery, Animesh Garg 외

Grounding textual phrases in visual content with standalone image-sentence pairs is a challenging task. When we consider grounding in instructional videos, this problem becomes profoundly more complex: the latent tempora…

Multiple Instance LearningSentenceVisual Grounding

Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding

2024-11-05 · Sombit Dey, Ozan Unal, Christos Sakaridis, Luc van Gool

3D visual grounding consists of identifying the instance in a 3D scene which is referred by an accompanying language description. While several architectures have been proposed within the commonly employed grounding-by-s…

3D visual groundingVisual Grounding

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

2026-07-09 · Haibin Tian, Huichao Xie, Xuelin Qian, Ruitao Lu 외 arxiv

Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) b…

Referring ExpressionVisual Grounding