paper-with-me

홈 › Papers

ReXamine-Global: A Framework for Uncovering Inconsistencies in Radiology Report Generation Metrics

2024-08-29 · Oishi Banerjee, Agustina Saenz, Kay Wu, Warren Clements, Adil Zia, Dominic Buensalido, Helen Kavnoudias, Alain S. Abi-Ghanem, Nour El Ghawi, Cibele Luna, Patricia Castillo, Khaled Al-Surimi, Rayyan A. Daghistani, Yuh-Min Chen, Heng-sheng Chao, Lars Heiliger, Moon Kim, Johannes Haubold, Frederic Jonske, Pranav Rajpurkar

Given the rapidly expanding capabilities of generative AI models for radiology, there is a need for robust metrics that can accurately measure the quality of AI-generated radiology reports across diverse hospitals. We develop ReXamine-Global, a LLM-powered, multi-site framework that tests metrics across different writing styles and patient populations, exposing gaps in their generalization. First, our method tests whether a metric is undesirably sensitive to reporting style, providing different scores depending on whether AI-generated reports are stylistically similar to ground-truth reports or not. Second, our method measures whether a metric reliably agrees with experts, or whether metric and expert scores of AI-generated report quality diverge for some sites. Using 240 reports from 6 hospitals around the world, we apply ReXamine-Global to 7 established report evaluation metrics and uncover serious gaps in their generalizability. Developers can apply ReXamine-Global when designing new report evaluation metrics, ensuring their robustness across sites. Additionally, our analysis of existing metrics can guide users of those metrics towards evaluation procedures that work reliably at their sites of interest.

📄 PDF Abstract BibTeX arXiv:2408.16208

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries

2025-05-13 · Zekun Wu, Seonglae Cho, Umar Mohammed, Cristian Munoz 외

Open-source AI libraries are foundational to modern AI systems but pose significant, underexamined risks across security, licensing, maintenance, supply chain integrity, and regulatory compliance. We present LibVulnWatch…

Exploring SAM Ablations for Enhancing Medical Segmentation in Radiology and Pathology

2023-09-30 · Amin Ranem, Niklas Babendererde, Moritz Fuchs, Anirban Mukhopadhyay

Medical imaging plays a critical role in the diagnosis and treatment planning of various medical conditions, with radiology and pathology heavily reliant on precise image segmentation. The Segment Anything Model (SAM) ha…

Brain Tumor SegmentationImage SegmentationSegmentationSemantic Segmentation+1

Uncovering Knowledge Gaps in Radiology Report Generation Models through Knowledge Graphs

2024-08-26 · Xiaoman Zhang, Julián N. Acosta, Hong-Yu Zhou, Pranav Rajpurkar

Recent advancements in artificial intelligence have significantly improved the automatic generation of radiology reports. However, existing evaluation methods fail to reveal the models' understanding of radiological imag…

Knowledge Graphs

RadDiff: Describing Differences in Radiology Image Sets with Natural Language

2026-01-07 · Xiaoxian Shen, Yuhui Zhang, Sahithi Ankireddy, Xiaohan Wang 외 arxiv

Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic system that performs radiologist-style co…

Natural Language UnderstandingMultimodal Reasoning

Uncovering Hidden Violent Tendencies in LLMs: A Demographic Analysis via Behavioral Vignettes

2025-06-25 · Quintin Myers, Yanjun Gao

Large language models (LLMs) are increasingly proposed for detecting and responding to violent content online, yet their ability to reason about morally ambiguous, real-world scenarios remains underexamined. We present t…

Text Generation