paper-with-me

Papers

StutterCut: Uncertainty-Guided Normalised Cut for Dysfluency Segmentation

2025-08-04 · Suhita Ghosh, Melanie Jouaiti, Jan-Ole Perschewski, Sebastian Stober arxiv

Detecting and segmenting dysfluencies is crucial for effective speech therapy and real-time feedback. However, most methods only classify dysfluencies at the utterance level. We introduce StutterCut, a semi-supervised framework that formulates dysfluency segmentation as a graph partitioning problem, where speech embeddings from overlapping windows are represented as graph nodes. We refine the connections between nodes using a pseudo-oracle classifier trained on weak (utterance-level) labels, with its influence controlled by an uncertainty measure from Monte Carlo dropout. Additionally, we extend the weakly labelled FluencyBank dataset by incorporating frame-level dysfluency boundaries for four dysfluency types. This provides a more realistic benchmark compared to synthetic datasets. Experiments on real and synthetic datasets show that StutterCut outperforms existing methods, achieving higher F1 scores and more precise stuttering onset detection.

📄 PDF Abstract BibTeX arXiv:2508.02255

Code (0)

등록된 구현이 없습니다.

Tasks

graph partitioning

Similar Papers 제목 키워드 기반

SSDM: Scalable Speech Dysfluency Modeling

2024-08-29 · Jiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk 외

Speech dysfluency modeling is the core module for spoken language learning, and speech therapy. However, there are three challenges. First, current state-of-the-art solutions\cite{lian2023unconstrained-udm, lian-anumanch…

Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection

2025-05-28 · Jinming Zhang, Xuanru Zhou, Jiachen Lian, Shuhe Li 외

Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled s…

DiversitySynthetic Data Generation

Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection

2025-05-22 · Chenxu Guo, Jiachen Lian, Xuanru Zhou, Jinming Zhang 외

Automatic detection of speech dysfluency aids speech-language pathologists in efficient transcription of disordered speech, enhancing diagnostics and treatment planning. Traditional methods, often limited to classificati…

Decoder

Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection

2024-09-15 · Xuanru Zhou, Cheol Jun Cho, Ayati Sharma, Brittany Morin 외

Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of tra…

object-detectionObject DetectionTemplate Matching

Unconstrained Dysfluency Modeling for Dysfluent Speech Transcription and Detection

2023-12-20 · Jiachen Lian, Carly Feng, Naasir Farooqi, Steve Li 외

Dysfluent speech modeling requires time-accurate and silence-aware transcription at both the word-level and phonetic-level. However, current research in dysfluency modeling primarily focuses on either transcription or de…