paper-with-me

Papers

Large Language Models for Dysfluency Detection in Stuttered Speech

2024-06-16 · Dominik Wagner, Sebastian P. Bayerl, Ilja Baumann, Korbinian Riedhammer, Elmar Nöth, Tobias Bocklet

Accurately detecting dysfluencies in spoken language can help to improve the performance of automatic speech and language processing components and support the development of more inclusive speech and language technologies. Inspired by the recent trend towards the deployment of large language models (LLMs) as universal learners and processors of non-lexical inputs, such as audio and video, we approach the task of multi-label dysfluency detection as a language modeling problem. We present hypotheses candidates generated with an automatic speech recognition system and acoustic representations extracted from an audio encoder model to an LLM, and finetune the system to predict dysfluency labels on three datasets containing English and German stuttered speech. The experimental results show that our system effectively combines acoustic and lexical information and achieves competitive results on the multi-label stuttering detection task.

📄 PDF Abstract BibTeX arXiv:2406.11025

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Dysfluencies Seldom Come Alone -- Detection as a Multi-Label Problem

2022-10-28 · Sebastian P. Bayerl, Dominik Wagner, Florian Hönig, Tobias Bocklet 외

Specially adapted speech recognition models are necessary to handle stuttered speech. For these to be used in a targeted manner, stuttered speech must be reliably detected. Recent works have treated stuttering as a multi…

Multi-class Classificationspeech-recognitionSpeech Recognition

A Stutter Seldom Comes Alone -- Cross-Corpus Stuttering Detection as a Multi-label Problem

2023-05-30 · Sebastian P. Bayerl, Dominik Wagner, Ilja Baumann, Florian Hönig 외

Most stuttering detection and classification research has viewed stuttering as a multi-class classification problem or a binary detection task for each dysfluency type; however, this does not match the nature of stutteri…

ClassificationCross-corpusMulti-class ClassificationMulti-Task Learning

On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts

2025-11-18 · Kashaf Gulzar, Dominik Wagner, Sebastian P. Bayerl, Florian Hönig 외 arxiv

Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting …

Speech Recognition

Deploying UDM Series in Real-Life Stuttered Speech Applications: A Clinical Evaluation Framework

2025-09-17 · Eric Zhang, Li Wei, Sarah Chen, Michael Wang arxiv

Stuttered and dysfluent speech detection systems have traditionally suffered from the trade-off between accuracy and clinical interpretability. While end-to-end deep learning models achieve high performance, their black-…

Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection

2025-05-28 · Jinming Zhang, Xuanru Zhou, Jiachen Lian, Shuhe Li 외

Speech dysfluency detection is crucial for clinical diagnosis and language assessment, but existing methods are limited by the scarcity of high-quality annotated data. Although recent advances in TTS model have enabled s…

DiversitySynthetic Data Generation