paper-with-me

홈 › Papers

End-to-End Speech Recognition and Disfluency Removal with Acoustic Language Model Pretraining

2023-09-08 · Saksham Bassi, Giulio Duregon, Siddhartha Jalagam, David Roth

The SOTA in transcription of disfluent and conversational speech has in recent years favored two-stage models, with separate transcription and cleaning stages. We believe that previous attempts at end-to-end disfluency removal have fallen short because of the representational advantage that large-scale language model pretraining has given to lexical models. Until recently, the high dimensionality and limited availability of large audio datasets inhibited the development of large-scale self-supervised pretraining objectives for learning effective audio representations, giving a relative advantage to the two-stage approach, which utilises pretrained representations for lexical tokens. In light of recent successes in large scale audio pretraining, we revisit the performance comparison between two-stage and end-to-end model and find that audio based language models pretrained using weak self-supervised objectives match or exceed the performance of similarly trained two-stage models, and further, that the choice of pretraining objective substantially effects a model's ability to be adapted to the disfluency removal task.

📄 PDF Abstract BibTeX arXiv:2309.04516

Code (1)

davidsroth/hubert-disfl 공식 구현 pytorch

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

End-to-End Speech Recognition and Disfluency Removal

2020-09-22 · Findings of the Association for Computational Linguistics 2020 · Paria Jamshid Lou, Mark Johnson

Disfluency detection is usually an intermediate step between an automatic speech recognition (ASR) system and a downstream task. By contrast, this paper aims to investigate the task of end-to-end speech recognition and d…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Multilingual Disfluency Removal using NMT

2016-12-01 · IWSLT 2016 12 · Eunah Cho, Jan Niehues, Thanh-Le Ha, Alex Waibel

In this paper, we investigate a multilingual approach for speech disfluency removal. A major challenge of this task comes from the costly nature of disfluency annotation. Motivated by the fact that speech disfluencies ar…

Machine TranslationNMTTranslation

Automatic Disfluency Detection from Untranscribed Speech

2023-11-01 · Amrit Romana, Kazuhito Koishida, Emily Mower Provost

Speech disfluencies, such as filled pauses or repetitions, are disruptions in the typical flow of speech. Stuttering is a speech disorder characterized by a high rate of disfluencies, but all individuals speak with some …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Natural Language Understandingspeech-recognition+1

Streaming Joint Speech Recognition and Disfluency Detection

2022-11-16 · Hayato Futami, Emiru Tsunoo, Kentaro Shibata, Yosuke Kashiwagi 외

Disfluency detection has mainly been solved in a pipeline approach, as post-processing of speech recognition. In this study, we propose Transformer-based encoder-decoder models that jointly solve speech recognition and d…

DecoderLanguage Modellingspeech-recognitionSpeech Recognition

Alzheimer's Dementia Recognition Using Acoustic, Lexical, Disfluency and Speech Pause Features Robust to Noisy Inputs

2021-06-29 · Morteza Rohanian, Julian Hough, Matthew Purver

We present two multimodal fusion-based deep learning models that consume ASR transcribed speech and acoustic data simultaneously to classify whether a speaker in a structured diagnostic task has Alzheimer's Disease and t…

Diagnostic