paper-with-me

홈 › Papers

PM-MMUT: Boosted Phone-Mask Data Augmentation using Multi-Modeling Unit Training for Phonetic-Reduction-Robust E2E Speech Recognition

2021-12-13 · Guodong Ma, Pengfei Hu, Nurmemet Yolwas, Shen Huang, Hao Huang

Consonant and vowel reduction are often encountered in speech, which might cause performance degradation in automatic speech recognition (ASR). Our recently proposed learning strategy based on masking, Phone Masking Training (PMT), alleviates the impact of such phenomenon in Uyghur ASR. Although PMT achieves remarkably improvements, there still exists room for further gains due to the granularity mismatch between the masking unit of PMT (phoneme) and the modeling unit (word-piece). To boost the performance of PMT, we propose multi-modeling unit training (MMUT) architecture fusion with PMT (PM-MMUT). The idea of MMUT framework is to split the Encoder into two parts including acoustic feature sequences to phoneme-level representation (AF-to-PLR) and phoneme-level representation to word-piece-level representation (PLR-to-WPLR). It allows AF-to-PLR to be optimized by an intermediate phoneme-based CTC loss to learn the rich phoneme-level context information brought by PMT. Experimental results on Uyghur ASR show that the proposed approaches outperform obviously the pure PMT. We also conduct experiments on the 960-hour Librispeech benchmark using ESPnet1, which achieves about 10% relative WER reduction on all the test set without LM fusion comparing with the latest official ESPnet1 pre-trained model.

📄 PDF Abstract BibTeX arXiv:2112.06721

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

CTC Loss 설명 없음

Similar Papers 제목 키워드 기반

Commuting to work and gender-conforming social norms: evidence from same-sex couples

2022-02-21 · Sonia Oreffice, Dario Sansone

We analyze work commute time by sexual orientation of partnered or married individuals, using the American Community Survey 2008-2019. Women in same-sex couples have a longer commute to work than working women in differe…

SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation

2022-11-02 · Yassine El Kheir, Shammur Absar Chowdhury, Ahmed Ali, Hamdy Mubarak 외

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronuncia…

Data AugmentationMulti-Task LearningPhone-level pronunciation scoring

Telephonetic: Making Neural Language Models Robust to ASR and Semantic Noise

2019-06-13 · Chris Larson, Tarek Lahlou, Diana Mingels, Zachary Kulis 외

Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Spe…

Data AugmentationDecoderLanguage ModelingLanguage Modelling+3

Through the Dual-Prism: A Spectral Perspective on Graph Data Augmentation for Graph Classification

2024-01-18 · Yutong Xia, Runpeng Yu, Yuxuan Liang, Xavier Bresson 외

Graph Neural Networks have become the preferred tool to process graph data, with their efficacy being boosted through graph data augmentation techniques. Despite the evolution of augmentation methods, issues like graph p…

Data AugmentationGraph Classification

Changes in Commuter Behavior from COVID-19 Lockdowns in the Atlanta Metropolitan Area

2023-02-27 · Tejas Santanam, Anthony Trasatti, Hanyu Zhang, Connor Riley 외

This paper analyzes the impact of COVID-19 related lockdowns in the Atlanta, Georgia metropolitan area by examining commuter patterns in three periods: prior to, during, and after the pandemic lockdown. A cellular phone …

ClusteringWord Embeddings