paper-with-me

홈 › Papers

FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores

2024-05-31 · Alyssa Huang, Oishi Banerjee, Kay Wu, Eduardo Pontes Reis, Pranav Rajpurkar

The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers of reports. In this work, we present FineRadScore, a Large Language Model (LLM)-based automated evaluation metric for generated CXR reports. Given a candidate report and a ground-truth report, FineRadScore gives the minimum number of line-by-line corrections required to go from the candidate to the ground-truth report. Additionally, FineRadScore provides an error severity rating with each correction and generates comments explaining why the correction was needed. We demonstrate that FineRadScore's corrections and error severity scores align with radiologist opinions. We also show that, when used to judge the quality of the report as a whole, FineRadScore aligns with radiologists as well as current state-of-the-art automated CXR evaluation metrics. Finally, we analyze FineRadScore's shortcomings to provide suggestions for future improvements.

📄 PDF Abstract BibTeX arXiv:2405.20613

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

VERT: Reliable LLM Judges for Radiology Report Evaluation

2026-04-03 · Federica Bologna, Jean-Philippe Corbeil, Matthew Wilkens, Asma Ben Abacha arxiv

Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear whether these approaches are robust when a…

parameter-efficient fine-tuning

QIAI at MEDIQA 2021: Multimodal Radiology Report Summarization

2021-06-01 · NAACL (BioNLP) 2021 6 · Jean-Benoit Delbrouck, Cassie Zhang, Daniel Rubin

This paper describes the solution of the QIAI lab sent to the Radiology Report Summarization (RRS) challenge at MEDIQA 2021. This paper aims to investigate whether using multimodality during training improves the summari…

A New Dataset for Summarizing Radiology Reports

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The radiology report summary is an important technology in smart healthcare. Compared with medical image processing and disease recognition which have been comprehensively studied, the research on radiology report summar…

Diagnostic

ReportQA: QA-Based Radiology Report Evaluation

2026-06-13 · Yiming Shi, Shaoshuai Yang, Xi Chen, Haolin Li 외 arxiv

Radiology report evaluation is essential for advancing automated report generation. Natural language generation metrics have limited clinical relevance. Clinical efficacy (CE) metrics evaluate important medical findings,…

Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation

2020-10-20 · NAACL 2021 4 · Yasuhide Miura, Yuhao Zhang, Emily Bao Tsai, Curtis P. Langlotz 외

Neural image-to-text radiology report generation systems offer the potential to improve radiology reporting by reducing the repetitive process of report drafting and identifying possible medical errors. However, existing…

Image to textNatural Language InferenceText Generation