paper-with-me

Papers

Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation

2023-10-05 · Kashu Yamazaki, Taisei Hanyu, Khoa Vo, Thang Pham, Minh Tran, Gianfranco Doretto, Anh Nguyen, Ngan Le

Precise 3D environmental mapping is pivotal in robotics. Existing methods often rely on predefined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, a groundbreaking approach for real-time open-vocabulary 3D mapping and queryable scene representation using RGB-D data. Open-Fusion harnesses the power of a pre-trained vision-language foundation model (VLFM) for open-set semantic comprehension and employs the Truncated Signed Distance Function (TSDF) for swift 3D scene reconstruction. By leveraging the VLFM, we extract region-based embeddings and their associated confidence maps. These are then integrated with 3D knowledge from TSDF using an enhanced Hungarian-based feature-matching mechanism. Notably, Open-Fusion delivers outstanding annotation-free 3D segmentation for open-vocabulary without necessitating additional 3D training. Benchmark tests on the ScanNet dataset against leading zero-shot methods highlight Open-Fusion's superiority. Furthermore, it seamlessly combines the strengths of region-based VLFM and TSDF, facilitating real-time 3D scene comprehension that includes object concepts and open-world semantics. We encourage the readers to view the demos on our project page: https://uark-aicv.github.io/OpenFusion

📄 PDF Abstract BibTeX arXiv:2310.03923

Code (1)

UARK-AICV/OpenFusion 공식 구현 pytorch

Tasks

3D Scene Reconstruction

Similar Papers 제목 키워드 기반

Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models

2023-03-08 · CVPR 2023 1 · Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon 외

We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation. Text-to-image diffusion …

Open Vocabulary Panoptic SegmentationOpen Vocabulary Semantic SegmentationOpen-World Instance SegmentationPanoptic Segmentation+3

OVI-MAP:Open-Vocabulary Instance-Semantic Mapping

2026-03-27 · Zilong Deng, Federico Tombari, Marc Pollefeys, Johanna Wald 외 arxiv

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, re…

Instance Segmentation

OpenMask3D: Open-Vocabulary 3D Instance Segmentation

2023-06-23 · NeurIPS 2023 11

We introduce the task of open-vocabulary 3D instance segmentation. Current approaches for 3D instance segmentation can typically only recognize object categories from a pre-defined closed set of classes that are annotate…

3D Instance Segmentation3D Open-Vocabulary Instance SegmentationInstance SegmentationObject+4

OpenHuman4D: Open-Vocabulary 4D Human Parsing

2025-07-14 · Keito Suzuki, Bang Du, Runfa Blark Li, Kunyao Chen 외 arxiv

Understanding dynamic 3D human representation has become increasingly critical in virtual and extended reality applications. However, existing human part segmentation methods are constrained by reliance on closed-set dat…

Human Part SegmentationVideo Object TrackingHuman Parsing

Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition

2026-09-15 · Yichi Zhang, Shenyue Wang, Jing Luo, Chunyang Yu 외 arxiv

Open-vocabulary multimodal emotion recognition (OV-MER) aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are dif…

Multimodal Emotion Recognition