paper-with-me

Papers

Multilingual Graphemic Hybrid ASR with Massive Data Augmentation

2019-09-14 · LREC 2020 5 · Chunxi Liu, Qiaochu Zhang, Xiaohui Zhang, Kritika Singh, Yatharth Saraf, Geoffrey Zweig

Towards developing high-performing ASR for low-resource languages, approaches to address the lack of resources are to make use of data from multiple languages, and to augment the training data by creating acoustic variations. In this work we present a single grapheme-based ASR model learned on 7 geographically proximal languages, using standard hybrid BLSTM-HMM acoustic models with lattice-free MMI objective. We build the single ASR grapheme set via taking the union over each language-specific grapheme set, and we find such multilingual graphemic hybrid ASR model can perform language-independent recognition on all 7 languages, and substantially outperform each monolingual ASR model. Secondly, we evaluate the efficacy of multiple data augmentation alternatives within language, as well as their complementarity with multilingual modeling. Overall, we show that the proposed multilingual graphemic hybrid ASR with various data augmentation can not only recognize any within training set languages, but also provide large ASR performance improvements.

📄 PDF Abstract BibTeX arXiv:1909.06522

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR

2019-10-22 · Duc Le, Thilo Koehler, Christian Fuegen, Michael L. Seltzer

Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, grap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

HIT-SCIR at MMNLU-22: Consistency Regularization for Multilingual Spoken Language Understanding

2023-01-05 · Bo Zheng, Zhouyang Li, Fuxuan Wei, Qiguang Chen 외

Multilingual spoken language understanding (SLU) consists of two sub-tasks, namely intent detection and slot filling. To improve the performance of these two sub-tasks, we propose to use consistency regularization based …

Data AugmentationIntent Detectionslot-fillingSlot Filling+1

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

2022-10-18 · Zhehuai Chen, Ankur Bapna, Andrew Rosenberg, Yu Zhang 외

Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model can be l…

Representation Learningspeech-recognitionSpeech RecognitionTransfer Learning

Deep Shallow Fusion for RNN-T Personalization

2020-11-16 · Duc Le, Gil Keren, Julian Chan, Jay Mahadeokar 외

End-to-end models in general, and Recurrent Neural Network Transducer (RNN-T) in particular, have gained significant traction in the automatic speech recognition community in the last few years due to their simplicity, c…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

2018-02-01 · Yu Wang, Xie Chen, Mark Gales, Anton Ragni 외

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition