paper-with-me

Papers

From Feedback to Checklists: Grounded Evaluation of AI-Generated Clinical Notes

2025-07-23 · Karen Zhou, John Giorgi, Pranav Mani, Peng Xu, Davis Liang, Chenhao Tan arxiv

AI-generated clinical notes are increasingly used in healthcare, but evaluating their quality remains a challenge due to high subjectivity and limited scalability of expert review. Existing automated metrics often fail to align with real-world physician preferences. To address this, we propose a pipeline that systematically distills real user feedback into structured checklists for note evaluation. These checklists are designed to be interpretable, grounded in human feedback, and enforceable by LLM-based evaluators. Using deidentified data from over 21,000 clinical encounters (prepared in accordance with the HIPAA safe harbor standard) from a deployed AI medical scribe system, we show that our feedback-derived checklist outperforms a baseline approach in our offline evaluations in coverage, diversity, and predictive power for human ratings. Extensive experiments confirm the checklist's robustness to quality-degrading perturbations, significant alignment with clinician preferences, and practical value as an evaluation methodology. In offline research settings, our checklist offers a practical tool for flagging notes that may fall short of our defined quality standards.

📄 PDF Abstract BibTeX arXiv:2507.17717

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation

2025-06-04 · Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha 외

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehe…

Multiple-choice

Learning Optimal Predictive Checklists

2021-12-02 · NeurIPS 2021 12 · Haoran Zhang, Quaid Morris, Berk Ustun, Marzyeh Ghassemi

Checklists are simple decision aids that are often used to promote safety and reliability in clinical applications. In this paper, we present a method to learn checklists for clinical decision support. We represent predi…

Fairness

Learning predictive checklists from continuous medical data

2022-11-14 · Yukti Makhija, Edward De Brouwer, Rahul G. Krishnan

Checklists, while being only recently introduced in the medical domain, have become highly popular in daily clinical practice due to their combined effectiveness and great interpretability. Checklists are usually designe…

TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

2024-10-04 · Jonathan Cook, Tim Rocktäschel, Jakob Foerster, Dennis Aumiller 외

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Preference judgments between model outputs hav…

AllInstruction Following

Consultation Checklists: Standardising the Human Evaluation of Medical Note Generation

2022-11-17 · Aleksandar Savkov, Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera 외

Evaluating automatically generated text is generally hard due to the inherently subjective nature of many aspects of the output quality. This difficulty is compounded in automatic consultation note generation by differin…