paper-with-me

Papers

ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery

2026-04-02 · Ke Li, Ting Wang, Di Wang, Yongshan Zhu, Yiming Zhang, Tao Lei, Quan Wang arxiv

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, which struggles to exploit fine-grained linguistic cues, such as \textit{spatial relations} and \textit{object attributes}, that are crucial for distinguishing objects with similar characteristics. Importantly, these cues play distinct roles across different grounding stages and should be leveraged accordingly to provide more explicit guidance. In this work, we propose \textbf{ProVG}, a novel RSVG framework that improves localization accuracy by decoupling language expressions into global context, spatial relations, and object attributes. To integrate these linguistic cues, ProVG employs a simple yet effective progressive cross-modal modulator, which dynamically modulates visual attention through a \textit{survey-locate-verify} scheme, enabling coarse-to-fine vision-language alignment. In addition, ProVG incorporates a cross-scale fusion module to mitigate the large-scale variations in remote sensing imagery, along with a language-guided calibration decoder to refine cross-modal alignment during prediction. A unified multi-task head further enables ProVG to support both referring expression comprehension and segmentation tasks. Extensive experiments on two benchmarks, \textit{i.e.}, RRSIS-D and RISBench, demonstrate that ProVG consistently outperforms existing methods, achieving new state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2604.01893

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionVisual Grounding

Similar Papers 제목 키워드 기반

Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding

2024-10-31 · Minghong Xie, Mengzhao Wang, Huafeng Li, Yafei Zhang 외

Visual grounding has attracted wide attention thanks to its broad application in various visual language tasks. Although visual grounding has made significant research progress, existing methods ignore the promotion effe…

ObjectPositionSentenceVisual Grounding

EDA: Explicit Text-Decoupling and Dense Alignment for 3D Visual Grounding

2022-09-29 · CVPR 2023 1 · Yanmin Wu, Xinhua Cheng, Renrui Zhang, Zesen Cheng 외

3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues. However, existing methods either extract the sentence-level features coupling …

3D visual groundingObjectSentenceVisual Grounding

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

2022-04-13 · CVPR 2022 1 · Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao 외

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detecti…

3D visual groundingVisual Grounding

Simulating Validity: Modal Decoupling in MLLM Generated Feedback on Science Drawings

2026-04-05 · Arne Bewersdorff, Nejla Yuruk, Xiaoming Zhai arxiv

In science education, students frequently construct hand-drawn visual models of scientific phenomena. These drawings rely on a visual structure where information is encoded through visual objects, their attributes, and r…

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

2026-05-20 · Anna Deichler, Jim O'Regan, Fethiye Irmak Dogan, Lubos Marcinek 외 arxiv

Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision-language models (VLMs) excel at static image tasks, they struggle to …

Visual Localization