paper-with-me

Papers

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models

2025-06-09 · Ruiyang Zhang, Hu Zhang, Hao Fei, Zhedong Zheng

Large Multimodal Models (LMMs), harnessing the complementarity among diverse modalities, are often considered more robust than pure Language Large Models (LLMs); yet do LMMs know what they do not know? There are three key open questions remaining: (1) how to evaluate the uncertainty of diverse LMMs in a unified manner, (2) how to prompt LMMs to show its uncertainty, and (3) how to quantify uncertainty for downstream tasks. In an attempt to address these challenges, we introduce Uncertainty-o: (1) a model-agnostic framework designed to reveal uncertainty in LMMs regardless of their modalities, architectures, or capabilities, (2) an empirical exploration of multimodal prompt perturbations to uncover LMM uncertainty, offering insights and findings, and (3) derive the formulation of multimodal semantic uncertainty, which enables quantifying uncertainty from multimodal responses. Experiments across 18 benchmarks spanning various modalities and 10 LMMs (both open- and closed-source) demonstrate the effectiveness of Uncertainty-o in reliably estimating LMM uncertainty, thereby enhancing downstream tasks such as hallucination detection, hallucination mitigation, and uncertainty-aware Chain-of-Thought reasoning.

📄 PDF Abstract BibTeX arXiv:2506.07575

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

Unveiling Deep Semantic Uncertainty Perception for Language-Anchored Multi-modal Vision-Brain Alignment

2025-11-06 · Zehui Feng, Chenqi Zhang, Mingru Wang, Minuo Wei 외 arxiv

Unveiling visual semantics from neural signals such as EEG, MEG, and fMRI remains a fundamental challenge due to subject variability and the entangled nature of visual features. Existing approaches primarily align neural…

Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models

2024-12-19 · Zijun Chen, WenBo Hu, Guande He, Zhijie Deng 외

Multimodal large language models (MLLMs) combine visual and textual data for tasks such as image captioning and visual question answering. Proper uncertainty calibration is crucial, yet challenging, for reliable use in a…

Autonomous DrivingImage CaptioningQuestion AnsweringVisual Question Answering

Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathways

2025-11-06 · Paloma Rabaey, Jong Hak Moon, Jung-Oh Lee, Min Gwan Kim 외 arxiv

Radiology reports are invaluable for clinical decision-making and hold great potential for automated analysis when structured into machine-readable formats. These reports often contain uncertainty, which we categorize in…

Image Classification

Uncertainty-Guided Latent Diagnostic Trajectory Learning for Sequential Clinical Diagnosis

2026-04-06 · Xuyang Shen, Haoran Liu, Dongjin Song, Martin Renqiang Min arxiv

Clinical diagnosis requires sequential evidence acquisition under uncertainty. However, most Large Language Model (LLM) based diagnostic systems assume fully observed patient information and therefore do not explicitly m…

Adversarial Vessel-Unveiling Semi-Supervised Segmentation for Retinopathy of Prematurity Diagnosis

2024-11-14 · Gozde Merve Demirci, Jiachen Yao, Ming-Chih Ho, Xiaoling Hu 외

Accurate segmentation of retinal images plays a crucial role in aiding ophthalmologists in diagnosing retinopathy of prematurity (ROP) and assessing its severity. However, due to their underdeveloped, thinner vessels, ma…

DiagnosticSegmentation