paper-with-me

홈 › Papers

Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding

2025-03-20 · CVPR 2025 1 · Jinlong Li, Cristiano Saltori, Fabio Poiesi, Nicu Sebe

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a common language space, which limits the potential of 3D models to leverage the diverse spatial and semantic capabilities encapsulated in various foundation models. In this paper, we propose Cross-modal and Uncertainty-aware Agglomeration for Open-vocabulary 3D Scene Understanding dubbed CUA-O3D, the first model to integrate multiple foundation models-such as CLIP, DINOv2, and Stable Diffusion-into 3D scene understanding. We further introduce a deterministic uncertainty estimation to adaptively distill and harmonize the heterogeneous 2D feature embeddings from these models. Our method addresses two key challenges: (1) incorporating semantic priors from VLMs alongside the geometric knowledge of spatially-aware vision foundation models, and (2) using a novel deterministic uncertainty estimation to capture model-specific uncertainties across diverse semantic and geometric sensitivities, helping to reconcile heterogeneous representations during training. Extensive experiments on ScanNetV2 and Matterport3D demonstrate that our method not only advances open-vocabulary segmentation but also achieves robust cross-domain alignment and competitive spatial perception capabilities. The code will be available at: https://github.com/TyroneLi/CUA_O3D.

📄 PDF Abstract BibTeX arXiv:2503.16707

Code (1)

tyroneli/cua_o3d 공식 구현 pytorch

Tasks

Scene Understanding

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

2025-06-09 · Ruiyang Zhang, Hu Zhang, Hao Fei, Zhedong Zheng

Large Multimodal Models (LMMs), harnessing the complementarity among diverse modalities, are often considered more robust than pure Language Large Models (LLMs); yet do LMMs know what they do not know? There are three ke…

Hallucination

Confidence-aware agglomeration classification and segmentation of 2D microscopic food crystal images

2025-07-31 · Xiaoyu Ji, Ali Shakouri, Fengqing Zhu arxiv

Food crystal agglomeration is a phenomenon occurs during crystallization which traps water between crystals and affects food product quality. Manual annotation of agglomeration in 2D microscopic images is particularly di…

Rank-Aware Agglomeration of Foundation Models for Immunohistochemistry Image Cell Counting

2025-11-16 · Zuqi Huang, Mengxin Tian, Huan Liu, Wentao Li 외 arxiv

Accurate cell counting in immunohistochemistry (IHC) images is critical for quantifying protein expression and aiding cancer diagnosis. However, the task remains challenging due to the chromogen overlap, variable biomark…

Uncertainty-Aware Boosted Ensembling in Multi-Modal Settings

2021-04-21 · Utkarsh Sarawgi, Rishab Khincha, Wazeer Zulfikar, Satrajit Ghosh 외

Reliability of machine learning (ML) systems is crucial in safety-critical applications such as healthcare, and uncertainty estimation is a widely researched method to highlight the confidence of ML systems in deployment…

Prediction Intervals

MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty

2024-01-23 · Tim Brödermann, David Bruggemann, Christos Sakaridis, Kevin Ta 외

Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic per…

Autonomous VehiclesObject DetectionPanoptic SegmentationSemantic Segmentation+1