paper-with-me

홈 › Papers

MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering

2025-09-15 · Wen-wai Yim, Asma Ben Abacha, Zixuan Yu, Robert Doerning, Fei Xia, Meliha Yetisgen arxiv

Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research.

📄 PDF Abstract BibTeX arXiv:2509.12405

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

2026-03-16 · Peigen Liu, Rui Ding, Yuren Mao, Ziyan Jiang 외 arxiv

Large Language Model (LLM)-based Collective Intelligence (CI) presents a promising approach to overcoming the data wall and continuously boosting the capabilities of LLM agents. However, there is currently no dedicated a…

MultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models

2024-02-14 · Corentin Royer, Bjoern Menze, Anjany Sekuboyina

We introduce MultiMedEval, an open-source toolkit for fair and reproducible evaluation of large, medical vision-language models (VLM). MultiMedEval comprehensively assesses the models' performance on a broad array of six…

BenchmarkingDiversity

MedPerf: Open Benchmarking Platform for Medical Artificial Intelligence using Federated Evaluation

2021-09-29 · Alexandros Karargyris, Renato Umeton, Micah J. Sheller, Alejandro Aristizabal 외

Medical AI has tremendous potential to advance healthcare by supporting the evidence-based practice of medicine, personalizing patient treatment, reducing costs, and improving provider and patient experience. We argue th…

BenchmarkingPhilosophy

OpenXAI: Towards a Transparent Evaluation of Model Explanations

2022-06-22 · Chirag Agarwal, Dan Ley, Satyapriya Krishna, Eshika Saxena 외

While several types of post hoc explanation methods have been proposed in recent literature, there is very little work on systematically benchmarking these methods. Here, we introduce OpenXAI, a comprehensive and extensi…

BenchmarkingExplainable Artificial Intelligence (XAI)Fairnessmodel

OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics

2025-06-14 · Vineeth Dorna, Anmol Mekala, Wenlong Zhao, Andrew McCallum 외

Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured. Yet the task is inherently challenging, partly d…

Benchmarking