paper-with-me

홈 › Papers

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

2024-12-05 · CVPR 2025 1 · Rong Li, Shijie Li, Lingdong Kong, Xulei Yang, Junwei Liang

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome these limitations, we introduce SeeGround, a zero-shot 3DVG framework leveraging 2D Vision-Language Models (VLMs) trained on large-scale 2D data. SeeGround represents 3D scenes as a hybrid of query-aligned rendered images and spatially enriched text descriptions, bridging the gap between 3D data and 2D-VLMs input formats. We propose two modules: the Perspective Adaptation Module, which dynamically selects viewpoints for query-relevant image rendering, and the Fusion Alignment Module, which integrates 2D images with 3D spatial descriptions to enhance object localization. Extensive experiments on ScanRefer and Nr3D demonstrate that our approach outperforms existing zero-shot methods by large margins. Notably, we exceed weakly supervised methods and rival some fully supervised ones, outperforming previous SOTA by 7.7% on ScanRefer and 7.1% on Nr3D, showcasing its effectiveness in complex 3DVG tasks.

📄 PDF Abstract BibTeX arXiv:2412.04383

Code (0)

등록된 구현이 없습니다.

Tasks

3D visual groundingObject LocalizationVisual Grounding

Similar Papers 제목 키워드 기반

Zero-Shot 3D Visual Grounding from Vision-Language Models

2025-05-28 · Rong Li, Shijie Li, Lingdong Kong, Xulei Yang 외

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on l…

3D visual groundingVisual Grounding

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

2026-05-25 · Cuong Huynh, Maxim Popov, Denis Gridusov, Sergey Kolyubin arxiv

3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language descriptions. Recent zero-shot methods leverage 2D vision-language models…

Visual GroundingPoint Clouds

Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding

2023-11-26 · CVPR 2024 1 · Zhihao Yuan, Jinke Ren, Chun-Mei Feng, Hengshuang Zhao 외

3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary, which can be restrictiv…

3D visual groundingObjectVisual Grounding

GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and Open-Vocabulary Object Detection

2023-12-22 · Haozhan Shen, Tiancheng Zhao, Mingwei Zhu, Jianwei Yin

Visual grounding, a crucial vision-language task involving the understanding of the visual context based on the query expression, necessitates the model to capture the interactions between objects, as well as various spa…

Attributeobject-detectionObject DetectionOpen-vocabulary object detection+2

Language-conditioned Detection Transformer

2023-11-29 · CVPR 2024 1 · Jang Hyun Cho, Philipp Krähenbühl

We present a new open-vocabulary detection framework. Our framework uses both image-level labels and detailed detection annotations when available. Our framework proceeds in three steps. We first train a language-conditi…

Pseudo Label