paper-with-me

Papers

Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis

2024-10-10 · Tuan Nguyen, Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard

With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pre-trained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.

📄 PDF Abstract BibTeX arXiv:2410.08250

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Temporally Explainable Dysarthric Speech Clarity Assessment

2025-05-31 · Seohyun Park, Chitralekha Gupta, Michelle Kah Yian Kwan, Xinhui Fung 외

Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric sp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Deep Learning for Pathological Speech: A Survey

2025-01-07 · Shakeel A. Sheikh, Md. Sahidullah, Ina Kodrasi

Advancements in spoken language technologies for neurodegenerative speech disorders are crucial for meeting both clinical and technological needs. This overview paper is vital for advancing the field, as it presents a co…

Automatic Speech RecognitionData AugmentationDeep Learningspeech-recognition+2

Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech

2024-09-13 · Pan-Pan Jiang, Jimmy Tobin, Katrin Tomanek, Robert L. MacDonald 외

Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

AI-Based Automated Speech Therapy Tools for persons with Speech Sound Disorders: A Systematic Literature Review

2022-04-21 · Chinmoy Deka, Abhishek Shrivastava, Ajish K. Abraham, Saurabh Nautiyal 외

This paper presents a systematic literature review of published studies on AI-based automated speech therapy tools for persons with speech sound disorders (SSD). The COVID-19 pandemic has initiated the requirement for au…

Systematic Literature Review

Towards Automated Assessment of Stuttering and Stuttering Therapy

2020-06-16 · Sebastian P. Bayerl, Florian Hönig, Joelle Reister, Korbinian Riedhammer

Stuttering is a complex speech disorder that can be identified by repetitions, prolongations of sounds, syllables or words, and blocks while speaking. Severity assessment is usually done by a speech therapist. While atte…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Positionspeech-recognition+1