paper-with-me

Papers

AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning

2026-01-23 · Suzhong Fu, Jingqi Dong, Xuan Ding, Rui Sun, Yiming Yang, Shuguang Cui, Zhen Li arxiv

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured diagnostic logic that underlies radiological interpretation, resulting in unreliable judgments and limited clinical relevance. We introduce AgentsEval, a multi-agent stream reasoning framework that emulates the collaborative diagnostic workflow of radiologists. By dividing the evaluation process into interpretable steps including criteria definition, evidence extraction, alignment, and consistency scoring, AgentsEval provides explicit reasoning traces and structured clinical feedback. We also construct a multi-domain perturbation-based benchmark covering five medical report datasets with diverse imaging modalities and controlled semantic variations. Experimental results demonstrate that AgentsEval delivers clinically aligned, semantically faithful, and interpretable evaluations that remain robust under paraphrastic, semantic, and stylistic perturbations. This framework represents a step toward transparent and clinically grounded assessment of medical report generation systems, fostering trustworthy integration of large language models into clinical practice.

📄 PDF Abstract BibTeX arXiv:2601.16685

Code (0)

등록된 구현이 없습니다.

Tasks

Medical Report Generation

Similar Papers 제목 키워드 기반

Explanation-Aware Learning for Enhanced Interpretability in Biomedical Imaging

2026-05-11 · Zubair Faruqui, Rahul Dubey arxiv

Deep neural networks for medical image diagnosis often achieve high predictive accuracy while relying on spurious or clinically irrelevant visual cues, limiting their trustworthiness in practice. Post-hoc explanation met…

Evaluating the Explainability of Vision Transformers in Medical Imaging

2025-10-13 · Leili Barekatain, Ben Glocker arxiv

Understanding model decisions is crucial in medical imaging, where interpretability directly impacts clinical trust and adoption. Vision Transformers (ViTs) have demonstrated state-of-the-art performance in diagnostic im…

Image Classification

XAI-MeD: Explainable Knowledge Guided Neuro-Symbolic Framework for Domain Generalization and Rare Class Detection in Medical Imaging

2026-01-05 · Midhat Urooj, Ayan Banerjee, Sandeep Gupta arxiv

Explainability domain generalization and rare class reliability are critical challenges in medical AI where deep models often fail under real world distribution shifts and exhibit bias against infrequent clinical conditi…

Diabetic Retinopathy GradingDomain Generalization

On the notion of missingness for path attribution explainability methods in medical settings: Guiding the selection of medically meaningful baselines

2025-08-20 · Alexander Geiger, Lars Wagner, Daniel Rueckert, Dirk Wilhelm 외 arxiv

The explainability of deep learning models remains a significant challenge, particularly in the medical domain where interpretable outputs are essential for clinical trust and transparency. Path attribution methods such …

Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging

2019-09-27 · Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, Christopher Ré

Machine learning models for medical image analysis often suffer from poor performance on important subsets of a population that are not identified during training or testing. For example, overall performance of a cancer …

BIG-bench Machine LearningMedical Image Analysis