paper-with-me

Papers

Recent Advances in Multi-modal 3D Scene Understanding: A Comprehensive Survey and Evaluation

2023-10-24 · Yinjie Lei, Zixuan Wang, Feng Chen, Guoqing Wang, Peng Wang, Yang Yang

Multi-modal 3D scene understanding has gained considerable attention due to its wide applications in many areas, such as autonomous driving and human-computer interaction. Compared to conventional single-modal 3D understanding, introducing an additional modality not only elevates the richness and precision of scene interpretation but also ensures a more robust and resilient understanding. This becomes especially crucial in varied and challenging environments where solely relying on 3D data might be inadequate. While there has been a surge in the development of multi-modal 3D methods over past three years, especially those integrating multi-camera images (3D+2D) and textual descriptions (3D+language), a comprehensive and in-depth review is notably absent. In this article, we present a systematic survey of recent progress to bridge this gap. We begin by briefly introducing a background that formally defines various 3D multi-modal tasks and summarizes their inherent challenges. After that, we present a novel taxonomy that delivers a thorough categorization of existing methods according to modalities and tasks, exploring their respective strengths and limitations. Furthermore, comparative results of recent approaches on several benchmark datasets, together with insightful analysis, are offered. Finally, we discuss the unresolved issues and provide several potential avenues for future research.

📄 PDF Abstract BibTeX arXiv:2310.15676

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingScene Understanding

Similar Papers 제목 키워드 기반

Deep Neural Networks for Visual Reasoning

2022-09-24 · Thao Minh Le

Visual perception and language understanding are - fundamental components of human intelligence, enabling them to understand and reason about objects and their interactions. It is crucial for machines to have this capaci…

Multimodal ReasoningVisual Reasoning

About Time: Advances, Challenges, and Outlooks of Action Understanding

2024-11-22 · Alexandros Stergiou, Ronald Poppe

We have witnessed impressive advances in video action understanding. Increased dataset sizes, variability, and computation availability have enabled leaps in performance and task diversification. Current systems can prov…

Action UnderstandingSurvey

Multimodal Language Models Cannot Spot Spatial Inconsistencies

2026-04-01 · Om Khangaonkar, Hadi J. Rad, Hamed Pirsiavash arxiv

Spatial consistency is a fundamental property of the visual world and a key requirement for models that aim to understand physical reality. Despite recent advances, multimodal large language models (MLLMs) often struggle…

What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?

2025-03-20 · Xuanming Cui, Jaiminkumar Ashokbhai Bhoi, Chionh Wei Peng, Adriel Kuek 외

Dynamic Scene Graph Generation (DSGG) for videos is a challenging task in computer vision. While existing approaches often focus on sophisticated architectural design and solely use recall during evaluation, we take a cl…

DecoderGraph GenerationScene Graph GenerationTriplet+1

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

2025-11-28 · Chen Li, Eric Peh, Basura Fernando arxiv

Recent advances in large vision-language models (VLMs) have shown significant promise for 3D scene understanding. Existing VLM-based approaches typically align 3D scene features with the VLM's embedding space. However, t…

Scene Understanding