paper-with-me

홈 › Papers

Voxel-informed Language Grounding

2022-05-19 · ACL 2022 5 · Rodolfo Corona, Shizhan Zhu, Dan Klein, Trevor Darrell

Natural language applied to natural 2D images describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3D geometric information in the form of voxel maps derived from the visual input using a volumetric reconstruction model. We show that VLG significantly improves grounding accuracy on SNARE, an object reference game task. At the time of writing, VLG holds the top place on the SNARE leaderboard, achieving SOTA results with a 2.0% absolute improvement.

📄 PDF Abstract BibTeX arXiv:2205.09710

Code (2)

rcorona/voxel_informed_language_grounding 공식 구현 pytorch
snaredataset/snare 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Voxel-informed Language Grounding

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Even when applied to 2D images, natural language describes a fundamentally 3D world. We present the Voxel-informed Language Grounder (VLG), a language grounding model that leverages 3D geometric information in the form …

A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding

2025-08-02 · Zhan Shi, Song Wang, Junbo Chen, Jianke Zhu arxiv

Visual grounding aims to identify objects or regions in a scene based on natural language descriptions, essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically d…

Autonomous DrivingDepth EstimationVisual Grounding

SeqAlign3DVG: A Sequence-Aligned Benchmark and Voxel Reasoning Framework for 3D Visual Grounding

2026-08-31 · Yi Zhang, Yi Wang, Yueting Wu, Kaiyue Yang 외 arxiv

Image-based 3D visual grounding is critical for embodied agents, yet existing benchmarks suffer from loose text-observation alignment and neglect temporal ordering. We introduce SeqAlign3DVG, a novel benchmark dedicated …

Visual GroundingPoint Clouds

Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial Attention

2024-10-29 · Haomeng Zhang, Chiao-An Yang, Raymond A. Yeh

Multi-object 3D Grounding involves locating 3D boxes based on a given query phrase from a point cloud. It is a challenging and significant task with numerous applications in visual understanding, human-computer interacti…

Object

Text-guided Sparse Voxel Pruning for Efficient 3D Visual Grounding

2025-02-14 · CVPR 2025 1 · Wenxuan Guo, Xiuwei Xu, Ziwei Wang, Jianjiang Feng 외

In this paper, we propose an efficient multi-level convolution architecture for 3D visual grounding. Conventional methods are difficult to meet the requirements of real-time inference due to the two-stage or point-based …

3D Object Detection3D visual groundingobject-detectionObject Detection+1