paper-with-me

홈 › Papers

Hallucination Benchmark in Medical Visual Question Answering

2024-01-11 · Jinge Wu, Yunsoo Kim, Honghan Wu

The recent success of large language and vision models (LLVMs) on vision question answering (VQA), particularly their applications in medicine (Med-VQA), has shown a great potential of realizing effective visual assistants for healthcare. However, these models are not extensively tested on the hallucination phenomenon in clinical settings. Here, we created a hallucination benchmark of medical images paired with question-answer sets and conducted a comprehensive evaluation of the state-of-the-art models. The study provides an in-depth analysis of current models' limitations and reveals the effectiveness of various prompting strategies.

📄 PDF Abstract BibTeX arXiv:2401.05827

Code (1)

knowlab/halt-medvqa 공식 구현 pytorch

Tasks

HallucinationMedical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding

2025-12-01 · Zahra Mahdavi, Zahra Khodakaramimaghsoud, Hooman Khaloo, Sina Bakhshandeh Taleshani 외 arxiv

Large vision-language models (LVLMs) are now central to healthcare applications such as medical visual question answering and imaging report generation. Yet, these models remain vulnerable to hallucination outputs that a…

Visual Question Answering

V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering

2026-01-26 · Mengyuan Jin, Zehui Liao, Yong Xia arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable capability in assisting disease diagnosis in medical visual question answering (VQA). However, their outputs remain vulnerable to hallucinations (i.e., respo…

Visual Question AnsweringComputational Efficiency

HALLUCINOGEN: A Benchmark for Evaluating Object Hallucination in Large Visual-Language Models

2024-12-29 · Ashish Seth, Dinesh Manocha, Chirag Agarwal

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance in performing complex multimodal tasks. However, they are still plagued by object hallucination: the misidentification or misclassification of…

HallucinationObjectObject HallucinationQuestion Answering+3

Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering

2025-03-26 · Zehui Liao, Shishuai Hu, Ke Zou, Huazhu Fu 외

Multimodal large language models (MLLMs) have demonstrated significant potential in medical Visual Question Answering (VQA). Yet, they remain prone to hallucinations-incorrect responses that contradict input images, posi…

DiagnosticHallucinationMedical Visual Question AnsweringQuestion Answering+2

VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering

2026-05-20 · Jiayi Chen, Benteng Ma, Zehui Liao, Winston Chong 외 arxiv

While medical Multimodal Large Language Models (MLLMs) have shown promise in assisting diagnosis, they still frequently generate hallucinated responses that appear linguistically plausible but lack visual evidence. Such …

Visual Question Answering