paper-with-me

홈 › Papers

MMScan: A Multi-Modal 3D Scene Dataset with Hierarchical Grounded Language Annotations

2024-06-13 · Ruiyuan Lyu, Jingli Lin, Tai Wang, Shuai Yang, Xiaohan Mao, Yilun Chen, Runsen Xu, Haifeng Huang, Chenming Zhu, Dahua Lin, Jiangmiao Pang

With the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress. However, limited by existing datasets, previous works mainly focus on understanding object properties or inter-object spatial relationships in a 3D scene. To tackle this problem, this paper builds the first largest ever multi-modal 3D scene dataset and benchmark with hierarchical grounded language annotations, MMScan. It is constructed based on a top-down logic, from region to object level, from a single target to inter-target relationships, covering holistic aspects of spatial and attribute understanding. The overall pipeline incorporates powerful VLMs via carefully designed prompts to initialize the annotations efficiently and further involve humans' correction in the loop to ensure the annotations are natural, correct, and comprehensive. Built upon existing 3D scanning data, the resulting multi-modal 3D dataset encompasses 1.4M meta-annotated captions on 109k objects and 7.7k regions as well as over 3.04M diverse samples for 3D visual grounding and question-answering benchmarks. We evaluate representative baselines on our benchmarks, analyze their capabilities in different aspects, and showcase the key problems to be addressed in the future. Furthermore, we use this high-quality dataset to train state-of-the-art 3D visual grounding and LLMs and obtain remarkable performance improvement both on existing benchmarks and in-the-wild evaluation. Codes, datasets, and benchmarks will be available at https://github.com/OpenRobotLab/EmbodiedScan.

📄 PDF Abstract BibTeX arXiv:2406.09401

Code (1)

openrobotlab/embodiedscan 공식 구현 pytorch

Tasks

3D visual groundingAttributeQuestion AnsweringVisual Grounding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MTANet: Multitask-Aware Network With Hierarchical Multimodal Fusion for RGB-T Urban Scene Understanding

2022-04-05 · journal 2022 4 · WuJie Zhou, Shaohua Dong, Jingsheng Lei, Lu Yu

Understanding urban scenes is a fundamental ability requirement for assisted driving and autonomous vehicles. Most of the available urban scene understanding methods use red-greenblue (RGB) images; however, their segme…

Autonomous VehiclesScene UnderstandingSegmentationThermal Image Segmentation

HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model

2025-11-28 · Chen Li, Eric Peh, Basura Fernando arxiv

Recent advances in large vision-language models (VLMs) have shown significant promise for 3D scene understanding. Existing VLM-based approaches typically align 3D scene features with the VLM's embedding space. However, t…

Scene Understanding

Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

2026-02-05 · Yiming Zhang, Qiangyu Yan, Borui Jiang, Kai Han arxiv

The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a "fast thinking" paradigm, reliant on end-to-end generation or e…

PRAM-R: A Perception-Reasoning-Action-Memory Framework with LLM-Guided Modality Routing for Adaptive Autonomous Driving

2026-03-04 · Yi Zhang, Xian Zhang, Saisi Zhao, Yinglei Song 외 arxiv

Multimodal perception enables robust autonomous driving but incurs unnecessary computational cost when all sensors remain active. This paper presents PRAM-R, a unified Perception-Reasoning-Action-Memory framework with LL…

Autonomous Driving

Cross-Modal and Hierarchical Modeling of Video and Text

2018-10-16 · ECCV 2018 9 · Bowen Zhang, Hexiang Hu, Fei Sha

Visual data and text data are composed of information at multiple granularities. A video can describe a complex scene that is composed of multiple clips or shots, where each depicts a semantically coherent event or actio…

Action RecognitionRetrievalTemporal Action LocalizationVideo Captioning+1