paper-with-me

홈 › Papers

VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation

2026-02-24 · Seongheon Park, Changdae Oh, Hyeong Kyu Choi, Sean Du, Sharon Li arxiv

Large Vision-Language Models (LVLMs) frequently hallucinate, limiting their safe deployment in real-world applications. Existing LLM self-evaluation methods rely on a model's ability to estimate the correctness of its own outputs, which can improve deployment reliability; however, they depend heavily on language priors and are therefore ill-suited for evaluating vision-conditioned predictions. We propose VAUQ, a vision-aware uncertainty quantification framework for LVLM self-evaluation that explicitly measures how strongly a model's output depends on visual evidence. VAUQ introduces the Image-Information Score (IS), which captures the reduction in predictive uncertainty attributable to visual input, and an unsupervised core-region masking strategy that amplifies the influence of salient regions. Combining predictive entropy with this core-masked IS yields a training-free scoring function that reliably reflects answer correctness. Comprehensive experiments show that VAUQ consistently outperforms existing self-evaluation methods across multiple datasets.

📄 PDF Abstract BibTeX arXiv:2602.21054

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Reference-free Hallucination Detection for Large Vision-Language Models

2024-08-11 · Qing Li, Jiahui Geng, Chenyang Lyu, Derui Zhu 외

Large vision-language models (LVLMs) have made significant progress in recent years. While LVLMs exhibit excellent ability in language understanding, question answering, and conversations of visual inputs, they are prone…

HallucinationQuestion AnsweringUncertainty Quantification

Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification

2026-02-05 · Tao Huang, Rui Wang, Xiaofei Liu, Yi Qin 외 arxiv

%Large vision-language models (LVLMs) have shown substantial advances in multimodal understanding and generation. However, when presented with incompetent or adversarial inputs, they frequently produce unreliable or even…

Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation

2026-05-26 · Joseph Hoche, David Brellmann, Gianni Franchi arxiv

Uncertainty quantification (UQ) remains a critical challenge in Large Vision Language Models (LVLMs) for reliable predictions and real-world deployment. However, most existing methods are adapted from the LLM literature …

Visual Grounding

Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation

2024-12-04 · CVPR 2025 1 · Gianni Franchi, Dat Nguyen Trong, Nacim Belkhir, Guoxuan Xia 외

Uncertainty quantification in text-to-image (T2I) generative models is crucial for understanding model behavior and improving output reliability. In this paper, we are the first to quantify and evaluate the uncertainty o…

Bias DetectionDisentanglementImage GenerationText to Image Generation+2

Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction

2025-04-24 · Yuanchang Ye, Weiyan Wen

This study addresses the critical challenge of hallucination mitigation in Large Vision-Language Models (LVLMs) for Visual Question Answering (VQA) tasks through a Split Conformal Prediction (SCP) framework. While LVLMs …

Conformal PredictionHallucinationPredictionQuestion Answering+3