paper-with-me

Papers

ReXTrust: A Model for Fine-Grained Hallucination Detection in AI-Generated Radiology Reports

2024-12-17 · Romain Hardy, Sung Eun Kim, Du Hyun Ro, Pranav Rajpurkar

The increasing adoption of AI-generated radiology reports necessitates robust methods for detecting hallucinations--false or unfounded statements that could impact patient care. We present ReXTrust, a novel framework for fine-grained hallucination detection in AI-generated radiology reports. Our approach leverages sequences of hidden states from large vision-language models to produce finding-level hallucination risk scores. We evaluate ReXTrust on a subset of the MIMIC-CXR dataset and demonstrate superior performance compared to existing approaches, achieving an AUROC of 0.8751 across all findings and 0.8963 on clinically significant findings. Our results show that white-box approaches leveraging model hidden states can provide reliable hallucination detection for medical AI systems, potentially improving the safety and reliability of automated radiology reporting.

📄 PDF Abstract BibTeX arXiv:2412.15264

Code (0)

등록된 구현이 없습니다.

Tasks

Hallucination

Similar Papers 제목 키워드 기반

A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy

2026-06-23 · Aminu Lawal, Niyoj Oli, Sachin Acharya, Prashnna Gyawali 외 arxiv

Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To date, most hallucination detection methods have been evaluated on radiology benchm…

Visual Question Answering

Fine-grained Hallucination Detection and Editing for Language Models

2024-01-12 · Abhika Mishra, Akari Asai, Vidhisha Balachandran, Yizhong Wang 외

Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in di…

HallucinationRetrieval

FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning

2024-10-08 · Ruosen Li, Ziming Luo, Xinya Du

Hallucinations in large language models (LLMs) pose significant challenges in tasks requiring complex multi-step reasoning, such as mathematical problem-solving. Existing approaches primarily detect the presence of hallu…

GSM8KHallucinationLanguage ModelingLanguage Modelling+3

Multilingual Fine-Grained News Headline Hallucination Detection

2024-07-22 · Jiaming Shen, Tianqi Liu, Jialu Liu, Zhen Qin 외

The popularity of automated news headline generation has surged with advancements in pre-trained language models. However, these models often suffer from the ``hallucination'' problem, where the generated headline is not…

HallucinationHeadline GenerationIn-Context Learning

Can Humans Dream of Electric Sheep? Human-Written Samples for Fine-Grained Vision-and-Language Hallucination Benchmarking

2026-08-02 · Timothee Mickus, Claudio Savelli, Eduardo Calò, Emilio Raimond 외 arxiv

In an age of rapid model turnover, how do we make hallucination evaluation more perennial? We explore whether human-written hallucination samples could take the place of model-generated hallucinations, in order to make b…