Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pre-trained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
Dysarthria, a motor speech disorder, affects intelligibility and requires targeted interventions for effective communication. In this work, we investigate automated mispronunciation feedback by collecting a dysarthric sp…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionDeep Learning for Pathological Speech: A Survey
Advancements in spoken language technologies for neurodegenerative speech disorders are crucial for meeting both clinical and technological needs. This overview paper is vital for advancing the field, as it presents a co…
Automatic Speech RecognitionData AugmentationDeep Learningspeech-recognition+2Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1AI-Based Automated Speech Therapy Tools for persons with Speech Sound Disorders: A Systematic Literature Review
This paper presents a systematic literature review of published studies on AI-based automated speech therapy tools for persons with speech sound disorders (SSD). The COVID-19 pandemic has initiated the requirement for au…
Systematic Literature ReviewTowards Automated Assessment of Stuttering and Stuttering Therapy
Stuttering is a complex speech disorder that can be identified by repetitions, prolongations of sounds, syllables or words, and blocks while speaking. Severity assessment is usually done by a speech therapist. While atte…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Positionspeech-recognition+1