paper-with-me

Papers

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

2026-06-15 · Sejal Bhalla, Larry Kieu, Aina Merchant, Eyal de Lara, Alex Mariakakis arxiv

Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems. Current clinical speech AI methods have largely progressed through isolated condition-specific studies, making results difficult to compare and generalization difficult to assess. We introduce SpeechDx, a large-scale benchmark for clinical speech AI spanning 12 datasets and 27 tasks across diverse health conditions. To enable evaluation across shared clinical mechanisms, SpeechDx structures tasks by the stage of speech production they disrupt: conceptualization, formulation, and articulation. The benchmark tests generalization by including tasks with limited labeled data and evaluating the same health condition across multiple datasets, distinguishing clinically meaningful patterns from dataset artefacts. We systematically evaluate 12 state-of-the-art audio encoders across all tasks and under zero-shot cross-condition transfer. Results show that large-scale speech models represent the strongest overall baselines, domain-specific models improve performance only on closely matched tasks, and no current representation generalizes reliably across the clinical speech landscape. SpeechDx establishes a shared evaluation framework for tracking progress toward general-purpose clinical speech representations

📄 PDF Abstract BibTeX arXiv:2606.17339

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Benchmark for Early-stage Parkinson's Disease Detection from Speech

2026-05-13 · Terry Yi Zhong, Cristian Tejedor-Garcia, Khiet P. Truong, Janna Maas 외 arxiv

Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard to compare because studies differ in datasets, languages, tasks, evaluation proto…

AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR

2023-09-30 · Tobi Olatunji, Tejumade Afonja, Aditya Yadavalli, Chris Chinenye Emezue 외

Africa has a very low doctor-to-patient ratio. At very busy clinics, doctors could see 30+ patients per day -- a heavy patient burden compared with developed countries -- but productivity tools such as clinical automatic…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment

2026-03-11 · Asif Azad, MD Sadik Hossain Shanto, Mohammad Sadat Hossain, Bdour Alwuqaysi 외 arxiv

Automated phoneme-level pronunciation assessment is vital for scalable speech therapy and language learning, yet validated tools for Arabic remain scarce. We present Harf-Speech, a modular system scoring Arabic pronuncia…

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

2026-04-29 · Darren Fürst, Sebastian Steindl, Ulrich Schäfer arxiv

Speech Sound Disorders (SSD) affect roughly five percent of children, yet speech-language pathologists face severe staffing shortages and unmanageable caseloads. We test a hierarchical approach to SSD classification on t…

Binary ClassificationSpeech RecognitionData Augmentation

Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces

2026-05-15 · Arne Nix, Robert James, Lasse Borgholt, Anna B. Ekner 외 arxiv

After decades of use in dictation and, more recently, ambient documentation, speech is emerging as a primary modality for interacting with technology and AI in healthcare. Yet medical speech recognition remains difficult…

Speech Recognition