paper-with-me

Papers

NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving

2025-03-28 · Fuhao Li, Huan Jin, Bin Gao, Liaoyuan Fan, Lihui Jiang, Long Zeng

Multi-view 3D visual grounding is critical for autonomous driving vehicles to interpret natural languages and localize target objects in complex environments. However, existing datasets and methods suffer from coarse-grained language instructions, and inadequate integration of 3D geometric reasoning with linguistic comprehension. To this end, we introduce NuGrounding, the first large-scale benchmark for multi-view 3D visual grounding in autonomous driving. We present a Hierarchy of Grounding (HoG) method to construct NuGrounding to generate hierarchical multi-level instructions, ensuring comprehensive coverage of human instruction patterns. To tackle this challenging dataset, we propose a novel paradigm that seamlessly combines instruction comprehension abilities of multi-modal LLMs (MLLMs) with precise localization abilities of specialist detection models. Our approach introduces two decoupled task tokens and a context query to aggregate 3D geometric information and semantic instructions, followed by a fusion decoder to refine spatial-semantic feature fusion for precise localization. Extensive experiments demonstrate that our method significantly outperforms the baselines adapted from representative 3D scene understanding methods by a significant margin and achieves 0.59 in precision and 0.64 in recall, with improvements of 50.8% and 54.7%.

📄 PDF Abstract BibTeX arXiv:2503.22436

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingAutonomous DrivingScene UnderstandingVisual Grounding

Similar Papers 제목 키워드 기반

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

2025-07-15 · Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu 외

3D visual grounding aims to identify and localize objects in a 3D space based on textual descriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolv…

3D visual groundingVisual Grounding

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding

2023-01-01 · ICCV 2023 1 · Zoey Guo, Yiwen Tang, Ray Zhang, Dong Wang 외

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality an…

3D visual groundingVisual Grounding

ViewRefer: Grasp the Multi-view Knowledge for 3D Visual Grounding with GPT and Prototype Guidance

2023-03-29 · Zoey Guo, Yiwen Tang, Ray Zhang, Dong Wang 외

Understanding 3D scenes from multi-view inputs has been proven to alleviate the view discrepancy issue in 3D visual grounding. However, existing methods normally neglect the view cues embedded in the text modality and fa…

3D visual groundingVisual Grounding

Multi-View Transformer for 3D Visual Grounding

2022-04-05 · CVPR 2022 1 · Shijia Huang, Yilun Chen, Jiaya Jia, LiWei Wang

The 3D visual grounding task aims to ground a natural language description to the targeted object in a 3D scene, which is usually represented in 3D point clouds. Previous works studied visual grounding under specific vie…

3D visual groundingVisual Grounding

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

2026-09-04 · Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu 외 arxiv

Recent zero-shot 3D visual grounding methods leverage vision-language models (VLMs) to localize objects in 3D scenes from natural language queries. However, these methods typically rely on heuristic rules to select which…

Natural Language QueriesVisual Grounding