paper-with-me

홈 › Papers

Multi-Object 3D Grounding with Dynamic Modules and Language-Informed Spatial Attention

2024-10-29 · Haomeng Zhang, Chiao-An Yang, Raymond A. Yeh

Multi-object 3D Grounding involves locating 3D boxes based on a given query phrase from a point cloud. It is a challenging and significant task with numerous applications in visual understanding, human-computer interaction, and robotics. To tackle this challenge, we introduce D-LISA, a two-stage approach incorporating three innovations. First, a dynamic vision module that enables a variable and learnable number of box proposals. Second, a dynamic camera positioning that extracts features for each proposal. Third, a language-informed spatial attention module that better reasons over the proposals to output the final prediction. Empirically, experiments show that our method outperforms the state-of-the-art methods on multi-object 3D grounding by 12.8% (absolute) and is competitive in single-object 3D grounding.

📄 PDF Abstract BibTeX arXiv:2410.22306

Code (1)

haomengz/D-LISA 공식 구현 pytorch

Tasks

Object

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Sigmoid Activation 설명 없음
Average Pooling 설명 없음

Similar Papers 제목 키워드 기반

3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point Clouds

2022-01-01 · CVPR 2022 1 · Daigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng 외

Observing that the 3D captioning task and the 3D grounding task contain both shared and complementary information in nature, in this work, we propose a unified framework to jointly solve these two distinct but closel…

3D dense captioningAttributeDense CaptioningVisual Grounding

Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding

2022-09-28 · Fengyuan Shi, Ruopeng Gao, Weilin Huang, LiMin Wang

Multimodal transformer exhibits high capacity and flexibility to align image and text for visual grounding. However, the existing encoder-only grounding framework (e.g., TransVG) suffers from heavy computation due to the…

DecoderVisual Grounding

NS3D: Neuro-Symbolic Grounding of 3D Objects and Relations

2023-03-23 · CVPR 2023 1 · Joy Hsu, Jiayuan Mao, Jiajun Wu

Grounding object properties and relations in 3D scenes is a prerequisite for a wide range of artificial intelligence tasks, such as visually grounded dialogues and embodied manipulation. However, the variability of the 3…

Question AnsweringReferring ExpressionReferring Expression ComprehensionVisual Reasoning

STORM: End-to-End Referring Multi-Object Tracking in Videos

2026-04-12 · Zijia Lu, Jingru Yi, Jue Wang, Yuxiao Chen 외 arxiv

Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding …

Multi-Object Tracking

Cross3DVG: Cross-Dataset 3D Visual Grounding on Different RGB-D Scans

2023-05-23 · Taiki Miyanishi, Daichi Azuma, Shuhei Kurita, Motoki Kawanabe

We present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent tendencie…

3D Reconstruction3D visual groundingVisual Grounding