paper-with-me

홈 › Papers

Frame-level SpecAugment for Deep Convolutional Neural Networks in Hybrid ASR Systems

2020-12-07 · Xinwei Li, Yuanyuan Zhang, Xiaodan Zhuang, Daben Liu

Inspired by SpecAugment -- a data augmentation method for end-to-end ASR systems, we propose a frame-level SpecAugment method (f-SpecAugment) to improve the performance of deep convolutional neural networks (CNN) for hybrid HMM based ASR systems. Similar to the utterance level SpecAugment, f-SpecAugment performs three transformations: time warping, frequency masking, and time masking. Instead of applying the transformations at the utterance level, f-SpecAugment applies them to each convolution window independently during training. We demonstrate that f-SpecAugment is more effective than the utterance level SpecAugment for deep CNN based hybrid models. We evaluate the proposed f-SpecAugment on 50-layer Self-Normalizing Deep CNN (SNDCNN) acoustic models trained with up to 25000 hours of training data. We observe f-SpecAugment reduces WER by 0.5-4.5% relatively across different ASR tasks for four languages. As the benefits of augmentation techniques tend to diminish as training data size increases, the large scale training reported is important in understanding the effectiveness of f-SpecAugment. Our experiments demonstrate that even with 25k training data, f-SpecAugment is still effective. We also demonstrate that f-SpecAugment has benefits approximately equivalent to doubling the amount of training data for deep CNNs.

📄 PDF Abstract BibTeX arXiv:2012.04094

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment

2020-04-02 · Wei Zhou, Wilfried Michel, Kazuki Irie, Markus Kitza 외

We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is successfully applied to improve performanc…

Data Augmentation

BCN2BRNO: ASR System Fusion for Albayzin 2020 Speech to Text Challenge

2021-01-29 · Martin Kocour, Guillermo Cámbara, Jordi Luque, David Bonet 외

This paper describes joint effort of BUT and Telef\'onica Research on development of Automatic Speech Recognition systems for Albayzin 2020 Challenge. We compare approaches based on either hybrid or end-to-end models. In…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

2019-04-18 · Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu 외

We present SpecAugment, a simple data augmentation method for speech recognition. SpecAugment is applied directly to the feature inputs of a neural network (i.e., filter bank coefficients). The augmentation policy consis…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+2

Transfer Learning and SpecAugment applied to SSVEP Based BCI Classification

2020-10-08 · Pedro R. A. S. Bassi, Willian Rampazzo, Romis Attux

Objective: We used deep convolutional neural networks (DCNNs) to classify electroencephalography (EEG) signals in a steady-state visually evoked potentials (SSVEP) based single-channel brain-computer interface (BCI), whi…

Brain Computer InterfaceClassificationData AugmentationEEG+6

Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems

2022-06-23 · Mingyu Cui, Jiajun Deng, Shoukang Hu, Xurong Xie 외

Fundamental modelling differences between hybrid and end-to-end (E2E) automatic speech recognition (ASR) systems create large diversity and complementarity among them. This paper investigates multi-pass rescoring and cro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1