paper-with-me

홈 › Papers

Towards Fluent Translations from Disfluent Speech

2018-11-07 · Elizabeth Salesky, Susanne Burger, Jan Niehues, Alex Waibel

When translating from speech, special consideration for conversational speech phenomena such as disfluencies is necessary. Most machine translation training data consists of well-formed written texts, causing issues when translating spontaneous speech. Previous work has introduced an intermediate step between speech recognition (ASR) and machine translation (MT) to remove disfluencies, making the data better-matched to typical translation text and significantly improving performance. However, with the rise of end-to-end speech translation systems, this intermediate step must be incorporated into the sequence-to-sequence architecture. Further, though translated speech datasets exist, they are typically news or rehearsed speech without many disfluencies (e.g. TED), or the disfluencies are translated into the references (e.g. Fisher). To generate clean translations from disfluent speech, cleaned references are necessary for evaluation. We introduce a corpus of cleaned target data for the Fisher Spanish-English dataset for this task. We compare how different architectures handle disfluencies and provide a baseline for removing disfluencies in end-to-end translation.

📄 PDF Abstract BibTeX arXiv:1811.03189

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Generating Fluent Translations from Disfluent Text Without Access to Fluent References: IIT Bombay@IWSLT2020

2020-07-01 · WS 2020 7 · Nikhil Saini, Jyotsana Khatri, Preethi Jyothi, Pushpak Bhattacharyya

Machine translation systems perform reasonably well when the input is well-formed speech or text. Conversational speech is spontaneous and inherently consists of many disfluencies. Producing fluent translations of disflu…

DenoisingMachine TranslationTranslation

Fluent Translations from Disfluent Speech in End-to-End Speech Translation

2019-06-03 · NAACL 2019 6 · Elizabeth Salesky, Matthias Sperber, Alex Waibel

Spoken language translation applications for speech suffer due to conversational speech phenomena, particularly the presence of disfluencies. With the rise of end-to-end speech translation models, processing steps such a…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Fluency Over Adequacy: A Pilot Study in Measuring User Trust in Imperfect MT

2018-02-16 · WS 2018 3 · Marianna J. Martindale, Marine Carpuat

Although measuring intrinsic quality has been a key factor in the advancement of Machine Translation (MT), successfully deploying MT requires considering not just intrinsic quality but also the user experience, including…

Machine TranslationTranslation

Inclusive ASR for Disfluent Speech: Cascaded Large-Scale Self-Supervised Learning with Targeted Fine-Tuning and Data Augmentation

2024-06-14 · Dena Mujtaba, Nihar R. Mahapatra, Megan Arney, J. Scott Yaruss 외

Automatic speech recognition (ASR) systems often falter while processing stuttering-related disfluencies -- such as involuntary blocks and word repetitions -- yielding inaccurate transcripts. A critical barrier to progre…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationSelf-Supervised Learning+2

Zero-shot Disfluency Detection for Indian Languages

2022-10-01 · COLING 2022 10 · Rohit Kundu, Preethi Jyothi, Pushpak Bhattacharyya

Disfluencies that appear in the transcriptions from automatic speech recognition systems tend to impair the performance of downstream NLP tasks. Disfluency correction models can help alleviate this problem. However, the …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition