paper-with-me

홈 › Papers

DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models

2026-01-02 · Yue Zhou, Jue Chen, Zilun Zhang, Penghui Huang, Ran Ding, Zhentao Zou, PengFei Gao, Yuchen Wei, Ke Li, Xue Yang, Xue Jiang, Hongxin Yang, Jonathan Li arxiv

Remote sensing (RS) large vision-language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explicit referring expressions-such as relative position, relative size, and color cues-thereby constraining performance on implicit VG tasks that require scenario-specific domain knowledge. This article introduces DVGBench, a high-quality implicit VG benchmark for drones, covering six major application scenarios: traffic, disaster, security, sport, social activity, and productive activity. Each object provides both explicit and implicit queries. Based on the dataset, we design DroneVG-R1, an LVLM that integrates the novel Implicit-to-Explicit Chain-of-Thought (I2E-CoT) within a reinforcement learning paradigm. This enables the model to take advantage of scene-specific expertise, converting implicit references into explicit ones and thus reducing grounding difficulty. Finally, an evaluation of mainstream models on both explicit and implicit VG tasks reveals substantial limitations in their reasoning capabilities. These findings provide actionable insights for advancing the reasoning capacity of LVLMs for drone-based agents. The code and datasets will be released at https://github.com/zytx121/DVGBench

📄 PDF Abstract BibTeX arXiv:2601.00998

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningVisual Grounding

Similar Papers 제목 키워드 기반

ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities

2024-07-01 · Chenming Zhu, Tai Wang, Wenwei Zhang, Kai Chen 외

Although great progress has been made in 3D visual grounding, current models still rely on explicit textual descriptions for grounding and lack the ability to reason human intentions from implicit instructions. We propos…

3D visual groundingLanguage ModelingLanguage ModellingLarge Language Model+1

Visual Intention Grounding for Egocentric Assistants

2025-04-18 · Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 외

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- …

ObjectVisual Grounding

Thinking with Visual Grounding

2026-06-15 · Junkai Zhang, Yihe Deng, Kai-Wei Chang, Wei Wang arxiv

Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces often leave the supporting image regions im…

Reinforcement LearningSpatial ReasoningVisual ReasoningVisual Grounding

Grounded 3D-Aware Spatial Vision-Language Modeling

2026-05-28 · An-Chieh Cheng, Yang Fu, Yatai Ji, Ligeng Zhu 외 arxiv

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introdu…

ChronusOmni: Improving Time Awareness of Omni Large Language Models

2025-12-10 · Yijing Chen, Yihan Wu, Kaisi Guan, Yuchen Ren 외 arxiv

Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on th…

Reinforcement Learning