paper-with-me

Papers

WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue

2025-11-20 · Zachary Ellis, Jared Joselowitz, Yash Deo, Yajie He, Anna Kalygina, Aisling Higham, Mana Rahimzadeh, Yan Jia, Ibrahim Habli, Ernest Lim arxiv

As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that standard, investigating whether WER or other common metrics correlate with the clinical impact of transcription errors. We establish a gold-standard benchmark by having expert clinicians compare ground-truth utterances to their ASR-generated counterparts, labeling the clinical impact of any discrepancies found in two distinct doctor-patient dialogue datasets. Our analysis reveals that WER and a comprehensive suite of existing metrics correlate poorly with the clinician-assigned risk labels (No, Minimal, or Significant Impact). To bridge this evaluation gap, we introduce an LLM-as-a-Judge, programmatically optimized using GEPA through DSPy to replicate expert clinical assessment. The optimized judge (Gemini-2.5-Pro) achieves human-comparable performance, obtaining 90% accuracy and a strong Cohen's kappa of 0.816. This work provides a validated, automated framework for moving ASR evaluation beyond simple textual fidelity to a necessary, scalable assessment of safety in clinical dialogue.

📄 PDF Abstract BibTeX arXiv:2511.16544

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Bridging Single Distortion Artifacts and Multifactorial Clinical Quality: Few-shot Biparametric MRI Quality Assessment via Distortion-trained Prototypical Networks

2026-06-17 · Yucheng Tang, Alexander Ng, Wen Yan, Natasha Thorley 외 arxiv

Clinical prostate multi-parametric MRI relies heavily on high-quality diffusion-weighted imaging (DWI), yet reading DWI is frequently compromised by geometric distortion, often caused by rectal air. Assessing quality via…

Image Quality AssessmentFew-Shot Learning

Viewport-Unaware Blind Omnidirectional Image Quality Assessment: A Flexible and Effective Paradigm

2025-03-08 · Jiebin Yan, Kangcheng Wu, Junjie Chen, Ziwen Tan 외

Most of existing blind omnidirectional image quality assessment (BOIQA) models rely on viewport generation by modeling user viewing behavior or transforming omnidirectional images (OIs) into varying formats; however, the…

ERPImage Quality Assessment

Potential of deep features for opinion-unaware, distortion-unaware, no-reference image quality assessment

2019-11-27 · Subhayan Mukherjee, Giuseppe Valenzise, Irene Cheng

Image Quality Assessment algorithms predict a quality score for a pristine or distorted input image, such that it correlates with human opinion. Traditional methods required a non-distorted "reference" version of the inp…

Image Quality AssessmentNo-Reference Image Quality Assessment

Learning to Distort: Weakly-Supervised Image Quality Transfer for Prostate DWI Correction

2026-06-17 · YuCheng Tang, Wen Yan, Alexander Ng, Natasha Thorley 외 arxiv

Single-shot echo-planar prostate diffusion-weighted imaging (DWI) is frequently complicated by geometric distortions, which impact the ability to derive reliable diagnoses from such images. Developing automated correctio…

Image Quality Assessment

A Feasibility Study of Answer-Unaware Question Generation for Education

2022-05-01 · Findings (ACL) 2022 5 · Liam Dugan, Eleni Miltsakaki, Shriyash Upadhyay, Etan Ginsberg 외

We conduct a feasibility study into the applicability of answer-unaware question generation models to textbook passages. We show that a significant portion of errors in such systems arise from asking irrelevant or un-int…

Question GenerationQuestion-Generation