paper-with-me

Papers

CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation

2025-05-22 · Yuyang Jiang, Chacha Chen, Shengyuan Wang, Feng Li, Zecong Tang, Benjamin M. Mervak, Lydia Chelala, Christopher M Straus, Reve Chahine, Samuel G. Armato III, Chenhao Tan

Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a Clinically-grounded tabular framework with Expert-curated labels and Attribute-level comparison for Radiology report evaluation (CLEAR). CLEAR not only examines whether a report can accurately identify the presence or absence of medical conditions, but also assesses whether it can precisely describe each positively identified condition across five key attributes: first occurrence, change, severity, descriptive location, and recommendation. Compared to prior works, CLEAR's multi-dimensional, attribute-level outputs enable a more comprehensive and clinically interpretable evaluation of report quality. Additionally, to measure the clinical alignment of CLEAR, we collaborate with five board-certified radiologists to develop CLEAR-Bench, a dataset of 100 chest X-ray reports from MIMIC-CXR, annotated across 6 curated attributes and 13 CheXpert conditions. Our experiments show that CLEAR achieves high accuracy in extracting clinical attributes and provides automated metrics that are strongly aligned with clinical judgment.

📄 PDF Abstract BibTeX arXiv:2505.16325

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDescriptive

Similar Papers 제목 키워드 기반

CCS: Clinical Consensus Selection for Radiology Report Generation

2026-05-28 · Xi Zhang, Yingshu Li, Zaiqiao Meng, Jake Lever 외 arxiv

Radiology report generation (RRG) is commonly formulated as a single-path generation task, where a multimodal large language model (MLLM) produces one decoded report as the final output. While recent progress has largely…

Decision Making

Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation

2025-08-04 · Radhika Dua, Young Joon, Kwon, Siddhant Dogra 외 arxiv

Radiological imaging is central to diagnosis, treatment planning, and clinical decision-making. Vision-language foundation models have spurred interest in automated radiology report generation (RRG), but safe deployment …

Question Answering

Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation

2025-12-18 · Sarosij Bose, Ravi K. Rajendran, Biplob Debnath, Konstantinos Karydis 외 arxiv

Radiology Report Generation (RRG) is a critical step toward automating healthcare workflows, facilitating accurate patient assessments, and reducing the workload of medical professionals. Despite recent progress in Large…

Medical Report GenerationVisual Reasoning

ReEvalMed: Rethinking Medical Report Evaluation by Aligning Metrics with Real-World Clinical Judgment

2025-09-30 · Ruochen Li, Jun Li, Bailiang Jian, Kun Yuan 외 arxiv

Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians' trust. This gap reveals fundamental flaws in how current metrics assess the quality of gen…

RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores

2025-08-21 · Yingshu Li, Yunyi Liu, Lingqiao Liu, Lei Wang 외 arxiv

Evaluating automatically generated radiology reports remains a fundamental challenge due to the lack of clinically grounded, interpretable, and fine-grained metrics. Existing methods either produce coarse overall scores …