paper-with-me

Papers

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

2026-06-12 · Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter arxiv

Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations. Prior work has focused on verbatim transcription and the integration of disfluency markers, but adapting models on limited datasets can lead to catastrophic forgetting of general-domain knowledge. We address this gap by leveraging continual learning (CL) with explicit disfluency tokens. We first introduce these tokens into a pretrained ASR model to establish stable token mechanisms, and then continue training on additional datasets with varying disfluency distributions. Through a detailed analysis of model dynamics during training, we identify a trade-off between marker learning and ASR performance, and a consistent cross-attention head mechanism shared across CL methods.

📄 PDF Abstract BibTeX arXiv:2606.14391

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningSpeech Recognition

Similar Papers 제목 키워드 기반

Incremental Disfluency Detection for Spoken Learner English

2022-07-01 · NAACL (BEA) 2022 7 · Lucy Skidmore, Roger Moore

Incremental disfluency detection provides a framework for computing communicative meaning from hesitations, repetitions and false starts commonly found in speech. One application of this area of research is in dialogue-b…

Adapting the NICT-JLE Corpus for Disfluency Detection Models

2023-08-04 · Lucy Skidmore, Roger K. Moore

The detection of disfluencies such as hesitations, repetitions and false starts commonly found in speech is a widely studied area of research. With a standardised process for evaluation using the Switchboard Corpus, mode…

TAG

Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis

2025-11-17 · Zaara Zabeen Arpa, Sadnam Sakib Apurbo, Nazia Karim Khan Oishee, Ajwad Abrar arxiv

Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitatio…

Speech Recognition

Typical vs. Atypical Disfluency Classification: Introducing the IIITH-TISA Corpus and Temporal Context-Based Feature Representations

2024-11-26 · Priyanka Kommagouni, Vamshiraghusimha Narasinga, Purva Barche, Sai Akarsh C 외

Speech disfluencies in spontaneous communication can be categorized as either typical or atypical. Typical disfluencies, such as hesitations and repetitions, are natural occurrences in everyday speech, while atypical dis…

Turn-Taking Prediction for Natural Conversational Speech

2022-08-29 · Shuo-Yiin Chang, Bo Li, Tara N. Sainath, Chao Zhang 외

While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency. Ho…

Predictionspeech-recognitionSpeech Recognition