paper-with-me

홈 › Papers

SpecAugment on Large Scale Datasets

2019-12-11 · Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V. Le, Yonghui Wu

Recently, SpecAugment, an augmentation scheme for automatic speech recognition that acts directly on the spectrogram of input utterances, has shown to be highly effective in enhancing the performance of end-to-end networks on public datasets. In this paper, we demonstrate its effectiveness on tasks with large scale datasets by investigating its application to the Google Multidomain Dataset (Narayanan et al., 2018). We achieve improvement across all test domains by mixing raw training data augmented with SpecAugment and noise-perturbed training data when training the acoustic model. We also introduce a modification of SpecAugment that adapts the time mask size and/or multiplicity depending on the length of the utterance, which can potentially benefit large scale tasks. By using adaptive masking, we are able to further improve the performance of the Listen, Attend and Spell model on LibriSpeech to 2.2% WER on test-clean and 5.2% WER on test-other.

📄 PDF Abstract BibTeX arXiv:1912.05533

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Frame-level SpecAugment for Deep Convolutional Neural Networks in Hybrid ASR Systems

2020-12-07 · Xinwei Li, Yuanyuan Zhang, Xiaodan Zhuang, Daben Liu

Inspired by SpecAugment -- a data augmentation method for end-to-end ASR systems, we propose a frame-level SpecAugment method (f-SpecAugment) to improve the performance of deep convolutional neural networks (CNN) for hyb…

Data Augmentation

RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification

2024-05-05 · June-Woo Kim, Miika Toikkanen, Sangmin Bae, Minseok Kim 외

Recent advancements in AI have democratized its deployment as a healthcare assistant. While pretrained models from large-scale visual and audio datasets have demonstrably generalized to this task, surprisingly, no studie…

Data AugmentationSound Classification

On Using SpecAugment for End-to-End Speech Translation

2019-11-20 · EMNLP (IWSLT) 2019 11 · Parnia Bahar, Albert Zeyer, Ralf Schlüter, Hermann Ney

This work investigates a simple data augmentation technique, SpecAugment, for end-to-end speech translation. SpecAugment is a low-cost implementation method applied directly to the audio input features and it consists of…

Data AugmentationTranslation

SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

2019-04-18 · Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu 외

We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consis…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+2

The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment

2020-04-02 · Wei Zhou, Wilfried Michel, Kazuki Irie, Markus Kitza 외

We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performanc…

Data Augmentation