paper-with-me

Papers

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models

2024-12-25 · Kaiwen Zuo, Yirui Jiang

Medical Large Language Models (MLLMs) have demonstrated potential in healthcare applications, yet their propensity for hallucinations -- generating medically implausible or inaccurate information -- presents substantial risks to patient care. This paper introduces MedHallBench, a comprehensive benchmark framework for evaluating and mitigating hallucinations in MLLMs. Our methodology integrates expert-validated medical case scenarios with established medical databases to create a robust evaluation dataset. The framework employs a sophisticated measurement system that combines automated ACHMI (Automatic Caption Hallucination Measurement in Medical Imaging) scoring with rigorous clinical expert evaluations and utilizes reinforcement learning methods to achieve automatic annotation. Through an optimized reinforcement learning from human feedback (RLHF) training pipeline specifically designed for medical applications, MedHallBench enables thorough evaluation of MLLMs across diverse clinical contexts while maintaining stringent accuracy standards. We conducted comparative experiments involving various models, utilizing the benchmark to establish a baseline for widely adopted large language models (LLMs). Our findings indicate that ACHMI provides a more nuanced understanding of the effects of hallucinations compared to traditional metrics, thereby highlighting its advantages in hallucination assessment. This research establishes a foundational framework for enhancing MLLMs' reliability in healthcare settings and presents actionable strategies for addressing the critical challenge of AI hallucinations in medical applications.

📄 PDF Abstract BibTeX arXiv:2412.18947

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucinationreinforcement-learningReinforcement Learning

Similar Papers 제목 키워드 기반

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

2025-02-28 · Qiao Yan, Yuchen Yuan, Xiaowei Hu, Yihan Wang 외

The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which the models may generate seemingly plausible results that are in fact incorrect.…

Decision MakingHallucinationQuestion AnsweringVisual Question Answering+1

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

2026-05-20 · Sunday Oyinlola Ogundoyin, Muhammad Ikram, Rahat Masood arxiv

Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to provide clinical guidance. However, they pose risks of hallucination, po…

Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

2024-06-14 · Jiawei Chen, Dingkang Yang, Tong Wu, Yue Jiang 외

Large Vision Language Models (LVLMs) are increasingly integral to healthcare applications, including medical visual question answering and imaging report generation. While these models inherit the robust capabilities of …

HallucinationMedical Visual Question AnsweringQuestion AnsweringVisual Question Answering

Toward Reliable Biomedical Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models

2025-05-20 · Guangzhi Xiong, Eric Xie, Corey Williams, Myles Kim 외

Large language models (LLMs) have shown significant potential in scientific disciplines such as biomedicine, particularly in hypothesis generation, where they can analyze vast literature, identify patterns, and suggest r…

Hallucinationscientific discovery

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

2026-03-05 · Qiao Jin, Yin Fang, Lauren He, Yifan Yang 외 arxiv

Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language models (LLMs) have the potential to automate this task, achieving strong performanc…