paper-with-me

Papers

Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment

2024-06-09 · Huma Ameer, Seemab Latif, Mehwish Fatima

The automated classification of stuttered speech has significant implications for timely assessments providing assistance to speech language pathologists. Despite notable advancements in the field, the cases in which multiple disfluencies occur in speech require attention. We have taken a progressive approach to fill this gap by classifying multi-stuttered speech more efficiently. The problem has been addressed by firstly curating a dataset of multi-stuttered disfluencies from open source dataset SEP-28k audio clips. Secondly, employing Whisper, a state-of-the-art speech recognition model has been leveraged by using its encoder and taking the problem as multi label classification. Thirdly, using a 6 encoder layer Whisper and experimenting with various layer freezing strategies, a computationally efficient configuration of the model was identified. The proposed configuration achieved micro, macro, and weighted F1-scores of 0.88, 0.85, and 0.87, correspondingly on an external test dataset i.e. Fluency-Bank. In addition, through layer freezing strategies, we were able to achieve the aforementioned results by fine-tuning a single encoder layer, consequently, reducing the model's trainable parameters from 20.27 million to 3.29 million. This research study unveils the contribution of the last encoder layer in the identification of disfluencies in stuttered speech. Consequently, it has led to a computationally efficient approach, 83.7% less parameters to train, making the proposed approach more adaptable for various dialects and languages.

📄 PDF Abstract BibTeX arXiv:2406.05784

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Dysfluencies Seldom Come Alone -- Detection as a Multi-Label Problem

2022-10-28 · Sebastian P. Bayerl, Dominik Wagner, Florian Hönig, Tobias Bocklet 외

Specially adapted speech recognition models are necessary to handle stuttered speech. For these to be used in a targeted manner, stuttered speech must be reliably detected. Recent works have treated stuttering as a multi…

Multi-class Classificationspeech-recognitionSpeech Recognition

Self-supervised Speech Models for Word-Level Stuttered Speech Detection

2024-09-16 · Yi-Jen Shih, Zoi Gkalitsiou, Alexandros G. Dimakis, David Harwath

Clinical diagnosis of stuttering requires an assessment by a licensed speech-language pathologist. However, this process is time-consuming and requires clinicians with training and experience in stuttering and fluency di…

Whisper in Focus: Enhancing Stuttered Speech Classification with Encoder Layer Optimization

2023-11-09 · Huma Ameer, Seemab Latif, Rabia Latif, Sana Mukhtar

In recent years, advancements in the field of speech processing have led to cutting-edge deep learning algorithms with immense potential for real-world applications. The automated identification of stuttered speech is on…

speech-recognitionSpeech Recognition

Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines

2026-04-22 · Hawau Olamide Toyin, Mutiah Apampa, Toluwani Aremu, Humaid Alblooshi 외 arxiv

Atypical speech is receiving greater attention in speech technology research, but much of this work unfolds with limited interdisciplinary dialogue. For stuttered speech in particular, it is widely recognised that curren…

Speech Recognition

AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection

2024-06-11 · Rong Gong, Hongfei Xue, Lezhi Wang, Xin Xu 외

The rapid advancements in speech technologies over the past two decades have led to human-level performance in tasks like automatic speech recognition (ASR) for fluent speech. However, the efficacy of these models dimini…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Event Detectionspeech-recognition+1