paper-with-me

홈 › Papers

3D Vision-Language Gaussian Splatting

2024-10-10 · Qucheng Peng, Benjamin Planche, Zhongpai Gao, Meng Zheng, Anwesa Choudhuri, Terrence Chen, Chen Chen, Ziyan Wu

Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D vision-language Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin.

📄 PDF Abstract BibTeX arXiv:2410.07577

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionAutonomous DrivingOpen Vocabulary Semantic SegmentationOpen-Vocabulary Semantic SegmentationRepresentation LearningScene UnderstandingSemantic Segmentation

Similar Papers 제목 키워드 기반

Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing

2024-04-01 · Ri-Zhao Qiu, Ge Yang, Weijia Zeng, Xiaolong Wang

Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however, demand the ability to manipulate both th…

Feature Splatting

ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation

2024-03-13 · Guanxing Lu, Shiyi Zhang, Ziwei Wang, Changliu Liu 외

Performing language-conditioned robotic manipulation tasks in unstructured environments is highly demanded for general intelligent robots. Conventional robotic manipulation methods usually learn semantic representation o…

Simulated Gaussian Manipulation

Occam's LGS: A Simple Approach for Language Gaussian Splatting

2024-12-02 · Jiahuan Cheng, Jan-Nico Zaech, Luc van Gool, Danda Pani Paudel

TL;DR: Gaussian Splatting is a widely adopted approach for 3D scene representation that offers efficient, high-quality 3D reconstruction and rendering. A major reason for the success of 3DGS is its simplicity of represen…

3DGS3D ReconstructionScene Understanding

AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration

2026-07-10 · Timofei Kozlov, Dmitrii Maliukov, Andrey Marchenko, Miguel Altamirano Cabrera 외 arxiv

We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, ou…

Pose Estimation

Cloth-Splatting: 3D Cloth State Estimation from RGB Supervision

2025-01-03 · Alberta Longhini, Marcel Büsching, Bardienus P. Duisterhof, Jens Lundell 외

We introduce Cloth-Splatting, a method for estimating 3D states of cloth from RGB images through a prediction-update framework. Cloth-Splatting leverages an action-conditioned dynamics model for predicting future states …

State Estimation