paper-with-me

Papers

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

2025-05-22 · Chenxu Guo, Jiachen Lian, Xuanru Zhou, Jinming Zhang, Shuhe Li, Zongli Ye, Hwi Joo Park, Anaisha Das, Zoe Ezzes, Jet Vonk, Brittany Morin, Rian Bogley, Lisa Wauters, Zachary Miller, Maria Gorno-Tempini, Gopala Anumanchipalli

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classification, provide insufficient clinical insight, and text-independent models misclassify dysfluency, especially in context-dependent cases. This work introduces Dysfluent-WFST, a zero-shot decoder that simultaneously transcribes phonemes and detects dysfluency. Unlike previous models, Dysfluent-WFST operates with upstream encoders like WavLM and requires no additional training. It achieves state-of-the-art performance in both phonetic error rate and dysfluency detection on simulated and real speech data. Our approach is lightweight, interpretable, and effective, demonstrating that explicit modeling of pronunciation behavior in decoding, rather than complex architectures, is key to improving dysfluency processing systems.

📄 PDF Abstract BibTeX arXiv:2505.16351

Code (1)

berkeley-speech-group/dysfluentwfst 공식 구현 pytorch

Tasks

Decoder

Similar Papers 제목 키워드 기반

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis

2025-06-05 · Zongli Ye, Jiachen Lian, Xuanru Zhou, Jinming Zhang 외

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting …

Integration of TensorFlow based Acoustic Model with Kaldi WFST Decoder

2019-06-21 · Minkyu Lim, Ji-Hwan Kim

While the Kaldi framework provides state-of-the-art components for speech recognition like feature extraction, deep neural network (DNN)-based acoustic models, and a weighted finite state transducer (WFST)-based decoder,…

Decoderspeech-recognitionSpeech Recognition

Unconstrained Dysfluency Modeling for Dysfluent Speech Transcription and Detection

2023-12-20 · Jiachen Lian, Carly Feng, Naasir Farooqi, Steve Li 외

Dysfluent speech modeling requires time-accurate and silence-aware transcription at both the word-level and phonetic-level. However, current research in dysfluency modeling primarily focuses on either transcription or de…

Arabic Code-Switching Speech Recognition using Monolingual Data

2021-07-04 · Ahmed Ali, Shammur Chowdhury, Amir Hussein, Yasser Hifny

Code-switching in automatic speech recognition (ASR) is an important challenge due to globalization. Recent research in multilingual ASR shows potential improvement over monolingual systems. We study key issues related t…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Differentiable Weighted Finite-State Transducers

2020-10-02 · Awni Hannun, Vineel Pratap, Jacob Kahn, Wei-Ning Hsu

We introduce a framework for automatic differentiation with weighted finite-state transducers (WFSTs) allowing them to be used dynamically at training time. Through the separation of graphs from operations on graphs, thi…

Handwriting Recognitionspeech-recognitionSpeech Recognition