paper-with-me

홈 › Papers

On Using SpecAugment for End-to-End Speech Translation

2019-11-20 · EMNLP (IWSLT) 2019 11 · Parnia Bahar, Albert Zeyer, Ralf Schlüter, Hermann Ney

This work investigates a simple data augmentation technique, SpecAugment, for end-to-end speech translation. SpecAugment is a low-cost implementation method applied directly to the audio input features and it consists of masking blocks of frequency channels, and/or time steps. We apply SpecAugment on end-to-end speech translation tasks and achieve up to +2.2\% \BLEU on LibriSpeech Audiobooks En->Fr and +1.2% on IWSLT TED-talks En->De by alleviating overfitting to some extent. We also examine the effectiveness of the method in a variety of data scenarios and show that the method also leads to significant improvements in various data conditions irrespective of the amount of training data.

📄 PDF Abstract BibTeX arXiv:1911.08876

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationTranslation

Similar Papers 제목 키워드 기반

SpliceOut: A Simple and Efficient Audio Augmentation Method

2021-09-30 · Arjit Jain, Pranay Reddy Samala, Deepak Mittal, Preethi Jyoti 외

Time masking has become a de facto augmentation technique for speech and audio tasks, including automatic speech recognition (ASR) and audio classification, most notably as a part of SpecAugment. In this work, we propose…

Audio ClassificationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Music Classification+4

Data Augmentation for End-to-End Speech Translation: FBK@IWSLT ‘19

2019-11-01 · EMNLP (IWSLT) 2019 11 · Mattia A. Di Gangi, Matteo Negri, Viet Nhat Nguyen, Amirhossein Tebbifakhr 외

This paper describes FBK’s submission to the end-to-end speech translation (ST) task at IWSLT 2019. The task consists in the “direct” translation (i.e. without intermediate discrete representation) of English speech data…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translation+4

Data Augmentation for End-to-end Code-switching Speech Recognition

2020-11-04 · Chenpeng Du, Hao Li, Yizhou Lu, Lan Wang 외

Training a code-switching end-to-end automatic speech recognition (ASR) model normally requires a large amount of data, while code-switching data is often limited. In this paper, three novel approaches are proposed for c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3

SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation

2020-02-27 · Arya D. McCarthy, Liezl Puzon, Juan Pino

We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio synthesized to resemble another speaker'…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationData Augmentation+4

SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

2019-04-18 · Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu 외

We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consis…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+2