paper-with-me

홈 › Papers

Harnessing Indirect Training Data for End-to-End Automatic Speech Translation: Tricks of the Trade

2019-09-14 · EMNLP (IWSLT) 2019 11 · Juan Pino, Liezl Puzon, Jiatao Gu, Xutai Ma, Arya D. McCarthy, Deepak Gopinath

For automatic speech translation (AST), end-to-end approaches are outperformed by cascaded models that transcribe with automatic speech recognition (ASR), then translate with machine translation (MT). A major cause of the performance gap is that, while existing AST corpora are small, massive datasets exist for both the ASR and MT subsystems. In this work, we evaluate several data augmentation and pretraining approaches for AST, by comparing all on the same datasets. Simple data augmentation by translating ASR transcripts proves most effective on the English--French augmented LibriSpeech dataset, closing the performance gap from 8.2 to 1.4 BLEU, compared to a very strong cascade that could directly utilize copious ASR and MT data. The same end-to-end approach plus fine-tuning closes the gap on the English--Romanian MuST-C dataset from 6.7 to 3.7 BLEU. In addition to these results, we present practical recommendations for augmentation and pretraining approaches. Finally, we decrease the performance gap to 0.01 BLEU using a Transformer-based architecture.

📄 PDF Abstract BibTeX arXiv:1909.06515

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationData AugmentationMachine Translationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Detecting Direct Speech in Multilingual Collection of 19th-century Novels

2020-05-01 · LREC 2020 5 · Joanna Byszuk, Micha{\l} Wo{\'z}niak, Mike Kestemont, Albert Le{\'s}niak 외

Fictional prose can be broadly divided into narrative and discursive forms with direct speech being central to any discourse representation (alongside indirect reported speech and free indirect discourse). This distincti…

Sentence

High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

2023-06-29 · Junchen Lu, Berrak Sisman, Mingyang Zhang, Haizhou Li

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the …

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Transformers in Speech Processing: A Survey

2023-03-21 · Siddique Latif, Aun Zaidi, Heriberto Cuayahuitl, Fahad Shamshad 외

The remarkable success of transformers in the field of natural language processing has sparked the interest of the speech-processing community, leading to an exploration of their potential for modeling long-range depende…

Automatic Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition+3

Toward a Web-based Speech Corpus for Algerian Dialectal Arabic Varieties

2017-04-01 · WS 2017 4 · Soumia Bougrine, Aicha Chorana, Abdallah Lakhdari, Hadda Cherroun

The success of machine learning for automatic speech processing has raised the need for large scale datasets. However, collecting such data is often a challenging task as it implies significant investment involving time …

Speech RecognitionSpeech Synthesis

Towards Accurate Text Verbalization for ASR Based on Audio Alignment

2019-09-01 · RANLP 2019 9 · Diana Geneva, Georgi Shopov

Verbalization of non-lexical linguistic units plays an important role in language modeling for automatic speech recognition systems. Most verbalization methods require valuable resources such as ground truth, large train…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2