paper-with-me

Papers

Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis

2025-06-05 · Zongli Ye, Jiachen Lian, Xuanru Zhou, Jinming Zhang, Haodong Li, Shuhe Li, Chenxu Guo, Anaisha Das, Peter Park, Zoe Ezzes, Jet Vonk, Brittany Morin, Rian Bogley, Lisa Wauters, Zachary Miller, Maria Gorno-Tempini, Gopala Anumanchipalli

Accurate alignment of dysfluent speech with intended text is crucial for automating the diagnosis of neurodegenerative speech disorders. Traditional methods often fail to model phoneme similarities effectively, limiting their performance. In this work, we propose Neural LCS, a novel approach for dysfluent text-text and speech-text alignment. Neural LCS addresses key challenges, including partial alignment and context-aware similarity mapping, by leveraging robust phoneme-level modeling. We evaluate our method on a large-scale simulated dataset, generated using advanced data simulation techniques, and real PPA data. Neural LCS significantly outperforms state-of-the-art models in both alignment accuracy and dysfluent speech segmentation. Our results demonstrate the potential of Neural LCS to enhance automated systems for diagnosing and analyzing speech disorders, offering a more accurate and linguistically grounded solution for dysfluent speech alignment.

📄 PDF Abstract BibTeX arXiv:2506.12073

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

2025-05-22 · Chenxu Guo, Jiachen Lian, Xuanru Zhou, Jinming Zhang 외

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classificati…

Decoder

YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection

2024-08-27 · Xuanru Zhou, Anshul Kashyap, Steve Li, Ayati Sharma 외

Dysfluent speech detection is the bottleneck for disordered speech analysis and spoken language learning. Current state-of-the-art models are governed by rule-based systems which lack efficiency and robustness, and are s…

Unconstrained Dysfluency Modeling for Dysfluent Speech Transcription and Detection

2023-12-20 · Jiachen Lian, Carly Feng, Naasir Farooqi, Steve Li 외

Dysfluent speech modeling requires time-accurate and silence-aware transcription at both the word-level and phonetic-level. However, current research in dysfluency modeling primarily focuses on either transcription or de…

Deploying UDM Series in Real-Life Stuttered Speech Applications: A Clinical Evaluation Framework

2025-09-17 · Eric Zhang, Li Wei, Sarah Chen, Michael Wang arxiv

Stuttered and dysfluent speech detection systems have traditionally suffered from the trade-off between accuracy and clinical interpretability. While end-to-end deep learning models achieve high performance, their black-…

Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning

2024-12-25 · Chirag Nagpal, Subhashini Venugopalan, Jimmy Tobin, Marilyn Ladewig 외

We introduce a large language model (LLM) capable of processing speech inputs and show that tuning it further with reinforcement learning on human preference (RLHF) enables it to adapt better to disordered speech than tr…

Language ModelingLanguage ModellingLarge Language Modelreinforcement-learning+3