paper-with-me

Papers

MUSE: Model-based Uncertainty-aware Similarity Estimation for zero-shot 2D Object Detection and Segmentation

2025-10-15 · Sungmin Cho, Sungbum Park, Insoo Oh arxiv

In this work, we introduce MUSE (Model-based Uncertainty-aware Similarity Estimation), a training-free framework designed for model-based zero-shot 2D object detection and segmentation. MUSE leverages 2D multi-view templates rendered from 3D unseen objects and 2D object proposals extracted from input query images. In the embedding stage, it integrates class and patch embeddings, where the patch embeddings are normalized using generalized mean pooling (GeM) to capture both global and local representations efficiently. During the matching stage, MUSE employs a joint similarity metric that combines absolute and relative similarity scores, enhancing the robustness of matching under challenging scenarios. Finally, the similarity score is refined through an uncertainty-aware object prior that adjusts for proposal reliability. Without any additional training or fine-tuning, MUSE achieves state-of-the-art performance on the BOP Challenge 2025, ranking first across the Classic Core, H3, and Industrial tracks. These results demonstrate that MUSE offers a powerful and generalizable framework for zero-shot 2D object detection and segmentation.

📄 PDF Abstract BibTeX arXiv:2510.17866

Code (0)

등록된 구현이 없습니다.

Tasks

2D Object Detection

Similar Papers 제목 키워드 기반

MUSE: Multimodal Uncertainty Quantification of State Estimation

2026-05-17 · Minkyung Kim, Henry Che, Bhargav Chandaka, Bhumsitt Pramuanpornsatid 외 arxiv

Accurate visual state estimation has been a central topic in robotics with a wide range of applications in robot navigation, autonomous driving, and autonomous flight. Recent advances in robot perception have led to sign…

Autonomous DrivingRobot Navigation

MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty

2024-01-23 · Tim Brödermann, David Bruggemann, Christos Sakaridis, Kevin Ta 외

Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic per…

Autonomous VehiclesObject DetectionPanoptic SegmentationSemantic Segmentation+1

ProMUSE: Progressive Multi-modal Uncertainty-guided Staged Evidential Alzheimer Disease Classification

2026-06-11 · Long Doan, Branden Chen, Ethan Litton, Huan Huang 외 arxiv

Alzheimer's disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population. Most treatments for AD are effective in the early stage, leading to an increasing demand for early AD diag…

Mind's Eye: Image Recognition by EEG via Multimodal Similarity-Keeping Contrastive Learning

2024-06-05 · Chi-Sheng Chen, Chun-Shu Wei

Decoding images from non-invasive electroencephalographic (EEG) signals has been a grand challenge in understanding how the human brain process visual information in real-world scenarios. To cope with the issues of signa…

Contrastive LearningEEGimage-classificationImage Classification+2

MUSEFood: Multi-sensor-based Food Volume Estimation on Smartphones

2019-03-18 · Junyi Gao, Weihao Tan, Liantao Ma, Yasha Wang 외

Researches have shown that diet recording can help people increase awareness of food intake and improve nutrition management, and thereby maintain a healthier life. Recently, researchers have been working on smartphone-b…

ManagementMulti-Task LearningNutrition