paper-with-me

Papers

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

2022-11-17 · Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Anya Belz, Ehud Reiter

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differing opinions between medical experts both about which patient statements should be included in generated notes and about their respective importance in arriving at a diagnosis. Previous real-world evaluations of note-generation systems saw substantial disagreement between expert evaluators. In this paper we propose a protocol that aims to increase objectivity by grounding evaluations in Consultation Checklists, which are created in a preliminary step and then used as a common point of reference during quality assessment. We observed good levels of inter-annotator agreement in a first evaluation study using the protocol; further, using Consultation Checklists produced in the study as reference for automatic metrics such as ROUGE or BERTScore improves their correlation with human judgements compared to using the original human note.

📄 PDF Abstract BibTeX arXiv:2211.09455

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models

2023-09-05 · Yusheng Liao, Yutong Meng, Hongcheng Liu, Yanfeng Wang 외

Large language models (LLMs) have achieved significant success in interacting with human. However, recent studies have revealed that these models often suffer from hallucinations, leading to overly confident but incorrec…

Multiple-choice

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

2025-07-23 · Karen Zhou, John Giorgi, Pranav Mani, Peng Xu 외 arxiv

AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review. Existing automated metrics often fail t…

LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation

2025-06-04 · Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha 외

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehe…

Multiple-choice

Learning predictive checklists from continuous medical data

2022-11-14 · Yukti Makhija, Edward De Brouwer, Rahul G. Krishnan

Checklists, while being only recently introduced in the medical domain, have become highly popular in daily clinical practice due to their combined effectiveness and great interpretability. Checklists are usually designe…

Multilingual CheckList: Generation and Evaluation

2022-03-24 · Karthikeyan K, Shaily Bhatt, Pankaj Singh, Somak Aditya 외

Multilingual evaluation benchmarks usually contain limited high-resource languages and do not test models for specific linguistic capabilities. CheckList is a template-based evaluation approach that tests models for spec…

DiversityMachine Translation