paper-with-me

홈 › Papers

RadSEM: A Finding-by-Finding Metric for Clinical Consistency in Radiology Reports

2026-06-03 · Zhenhong Yang, Zhuoyun Liu, Jintao Fei, Wen Tang, Shichao Quan, Jun Zhao, Jun Xu arxiv

Radiology report evaluation must distinguish clinical compatibility from surface similarity, because negation, laterality, or normal-abnormal polarity can reverse a finding. We propose RadSEM (Radiology Sentence-Level Evaluation Metric), a constrained LLM-assisted metric for reference-based evaluation of radiology Findings. RadSEM rewrites reference and generated reports into ordered atomic finding sentences, each expressing one site-finding proposition. It then performs contradiction-constrained many-to-many matching: incompatible pairs such as "effusion" and "no effusion" receive no credit, while compatible granularity differences can receive partial credit. A deterministic stage weights pairs by part-whole and abnormal-detail relationships, counts unmatched findings, and produces an abnormal-focused weighted F1 score. Thus, the LLM supports structured rewriting and local alignment rather than acting as an opaque judge. We evaluate RadSEM with SSREE, a controlled monotonicity stress test built from 2,448 de-identified reports expanded into five graded corruption levels. RadSEM achieves Kendall tau_b of 0.957, all-pairs concordance of 97.8%, adjacent concordance of 95.0%, and strict five-level ordering for 81.9% of reports, outperforming radiology-specific and general text metrics while avoiding the failure in which polarity-inverted reports regain lexical overlap. On the same SSREE set, RadSEM outperforms the Ref-anchored RadSEM-Alt policy, improving adjacent concordance from 90.7% to 95.0% and strict ordering from 67.2% to 81.9%. On a 599-triplet synonym/antonym subset, RadSEM prefers synonyms in 597 cases (99.67%). These results suggest that explicit finding units, contradiction-aware matching, and abnormal-focused deterministic scoring make report scoring more interpretable and sensitive to clinically meaningful errors. Code is available at https://github.com/jdh-algo/RadSEM.

📄 PDF Abstract BibTeX arXiv:2606.17062

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlexR: Few-shot Classification with Language Embeddings for Structured Reporting of Chest X-rays

2022-03-29 · Matthias Keicher, Kamilia Zaripova, Tobias Czempiel, Kristina Mach 외

The automation of chest X-ray reporting has garnered significant interest due to the time-consuming nature of the task. However, the clinical accuracy of free-text reports has proven challenging to quantify using natural…

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

2026-02-27 · Vikash Singh, Debargha Ganguly, Haotian Yu, Chengwei Zhou 외 arxiv

Vision-language models (VLMs) show promise in drafting radiology reports, yet they frequently suffer from logical inconsistencies, generating diagnostic impressions unsupported by their own perceptual findings or missing…

Clinical Knowledge

CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation

2026-04-27 · Ruifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi 외 arxiv

The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of findings, and the presence of fine-grained,…

Gender Bias in Large Language Models for Healthcare: Assignment Consistency and Clinical Implications

2025-10-08 · Mingxuan Liu, Yuhe Ke, Wentao Zhu, Mayli Mertens 외 arxiv

The integration of large language models (LLMs) into healthcare holds promise to enhance clinical decision-making, yet their susceptibility to biases remains a critical concern. Gender has long influenced physician behav…

Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation

2025-08-04 · Radhika Dua, Young Joon, Kwon, Siddhant Dogra 외 arxiv

Radiological imaging is central to diagnosis, treatment planning, and clinical decision-making. Vision-language foundation models have spurred interest in automated radiology report generation (RRG), but safe deployment …

Question Answering