paper-with-me

홈 › Papers

ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

2019-12-18 · ECCV 2020 8 · Dave Zhenyu Chen, Angel X. Chang, Matthias Nießner

We introduce the task of 3D object localization in RGB-D scans using natural language descriptions. As input, we assume a point cloud of a scanned 3D scene along with a free-form description of a specified target object. To address this task, we propose ScanRefer, learning a fused descriptor from 3D object proposals and encoded sentence embeddings. This fused descriptor correlates language expressions with geometric features, enabling regression of the 3D bounding box of a target object. We also introduce the ScanRefer dataset, containing 51,583 descriptions of 11,046 objects from 800 ScanNet scenes. ScanRefer is the first large-scale effort to perform object localization via natural language expression directly in 3D.

📄 PDF Abstract BibTeX arXiv:1912.08830

Code (3)

chodi150/scanrefer-pointgroup pytorch
daveredrum/ScanRefer pytorch
yuhaonankaka/adl_maskrcnn pytorch

Tasks

ObjectObject LocalizationregressionSentenceSentence Embeddings

Similar Papers 제목 키워드 기반

Scan2Cap: Context-aware Dense Captioning in RGB-D Scans

2020-12-03 · CVPR 2021 1 · Dave Zhenyu Chen, Ali Gholami, Matthias Nießner, Angel X. Chang

We introduce the task of dense captioning in 3D scans from commodity RGB-D sensors. As input, we assume a point cloud of a 3D scene; the expected output is the bounding boxes along with the descriptions for the underlyin…

3D dense captioning3D Object Detection3D Question Answering (3D-QA)Dense Captioning+4

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

2026-03-18 · Wenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie 외 arxiv

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achie…

Referring ExpressionVisual GroundingPoint Clouds

Multi3DRefer: Grounding Text Description to Multiple 3D Objects

2023-09-11 · ICCV 2023 1 · Yiming Zhang, ZeMing Gong, Angel X. Chang

We introduce the task of localizing a flexible number of objects in real-world 3D scenes using natural language descriptions. Existing 3D visual grounding tasks focus on localizing a unique object given a text descriptio…

3D visual groundingContrastive LearningObjectObject Rearrangement+3

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

2025-10-11 · Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang arxiv

Localizing 3D objects using natural language is essential for robotic scene understanding. The descriptions often involve multiple spatial relationships to distinguish similar objects, making 3D-language alignment diffic…

Scene Understanding

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

2024-12-05 · CVPR 2025 1 · Rong Li, Shijie Li, Lingdong Kong, Xulei Yang 외

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and …

3D visual groundingObject LocalizationVisual Grounding