paper-with-me

홈 › Papers

MASSTAR: A Multi-Modal and Large-Scale Scene Dataset with a Versatile Toolchain for Surface Prediction and Completion

2024-03-18 · Guiyong Zheng, Jinqi Jiang, Chen Feng, Shaojie Shen, Boyu Zhou

Surface prediction and completion have been widely studied in various applications. Recently, research in surface completion has evolved from small objects to complex large-scale scenes. As a result, researchers have begun increasing the volume of data and leveraging a greater variety of data modalities including rendered RGB images, descriptive texts, depth images, etc, to enhance algorithm performance. However, existing datasets suffer from a deficiency in the amounts of scene-level models along with the corresponding multi-modal information. Therefore, a method to scale the datasets and generate multi-modal information in them efficiently is essential. To bridge this research gap, we propose MASSTAR: a Multi-modal lArge-scale Scene dataset with a verSatile Toolchain for surfAce pRediction and completion. We develop a versatile and efficient toolchain for processing the raw 3D data from the environments. It screens out a set of fine-grained scene models and generates the corresponding multi-modal data. Utilizing the toolchain, we then generate an example dataset composed of over a thousand scene-level models with partial real-world data added. We compare MASSTAR with the existing datasets, which validates its superiority: the ability to efficiently extract high-quality models from complex scenarios to expand the dataset. Additionally, several representative surface completion algorithms are benchmarked on MASSTAR, which reveals that existing algorithms can hardly deal with scene-level completion. We will release the source code of our toolchain and the dataset. For more details, please see our project page at https://sysu-star.github.io/MASSTAR.

📄 PDF Abstract BibTeX arXiv:2403.11681

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Large-Scale Outdoor Multi-modal Dataset and Benchmark for Novel View Synthesis and Implicit Scene Reconstruction

2023-01-17 · ICCV 2023 1 · Chongshan Lu, Fukun Yin, Xin Chen, Tao Chen 외

Neural Radiance Fields (NeRF) has achieved impressive results in single object scene reconstruction and novel view synthesis, which have been demonstrated on many single modality and single object focused indoor scene da…

NeRFNovel View SynthesisSurface Reconstruction

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

2025-03-04 · CVPR 2025 1 · Ege Özsoy, Chantal Pellegrini, Tobias Czempiel, Felix Tristram 외

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient …

2D Panoptic SegmentationGraph GenerationLanguage ModelingLanguage Modelling+2

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

2025-07-17 · Penglei Sun, Yaoxian Song, Xiangru Zhu, Xiang Liu 외

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have primarily focused on indoor household tasks, …

Question AnsweringScene Understanding

VGStore: A Multimodal Extension to SPARQL for Querying RDF Scene Graph

2022-09-07 · Yanzeng Li, Zilong Zheng, Wenjuan Han, Lei Zou

Semantic Web technology has successfully facilitated many RDF models with rich data representation methods. It also has the potential ability to represent and store multimodal knowledge bases such as multimodal scene gra…

Relational ReasoningSemantic SimilaritySemantic Textual Similarity

Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction

2025-09-22 · Sitian Shen, Georgi Pramatarov, Yifu Tao, Daniele De Martini arxiv

This paper proposes Neural-MMGS, a novel neural 3DGS framework for multimodal large-scale scene reconstruction that fuses multiple sensing modalities in a per-gaussian compact, learnable embedding. While recent works foc…