Multi-pass Training and Cross-information Fusion for Low-resource End-to-end Accented Speech Recognition
Low-resource accented speech recognition is one of the important challenges faced by current ASR technology in practical applications. In this study, we propose a Conformer-based architecture, called Aformer, to leverage both the acoustic information from large non-accented and limited accented training data. Specifically, a general encoder and an accent encoder are designed in the Aformer to extract complementary acoustic information. Moreover, we propose to train the Aformer in a multi-pass manner, and investigate three cross-information fusion methods to effectively combine the information from both general and accent encoders. All experiments are conducted on both the accented English and Mandarin ASR tasks. Results show that our proposed methods outperform the strong Conformer baseline by relative 10.2% to 24.5% word/character error rate reduction on six in-domain and out-of-domain accented test sets.
Code (0)
등록된 구현이 없습니다.
Tasks
Accented Speech Recognitionspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Cross-Modal Message Passing for Two-stream Fusion
Processing and fusing information among multi-modal is a very useful technique for achieving high performance in many computer vision problems. In order to tackle multi-modal information more effectively, we introduce a …
Action RecognitionGeneral ClassificationOptical Flow EstimationTemporal Action Localization+1COMPASS: Complete Multimodal Fusion via Proxy Tokens and Shared Spaces for Ubiquitous Sensing
Missing modalities in multimodal sensing cause not only information loss but also a fusion-interface mismatch: a fusion head trained on a canonical set of modality slots must operate on changing observed subsets at infer…
Learning deep multiresolution representations for pansharpening
Retaining spatial characteristics of panchromatic image and spectral information of multispectral bands is a critical issue in pansharpening. This paper proposes a pyramid based deep fusion framework that preserves spect…
Pansharpening2DPASS: 2D Priors Assisted Semantic Segmentation on LiDAR Point Clouds
As camera and LiDAR sensors capture complementary information used in autonomous driving, great efforts have been made to develop semantic segmentation algorithms through multi-modality data fusion. However, fusion-based…
3D Semantic SegmentationAutonomous DrivingKnowledge DistillationLIDAR Semantic Segmentation+3Learning Fair Graph Representations with Multi-view Information Bottleneck
Graph neural networks (GNNs) excel on relational data by passing messages over node features and structure, but they can amplify training data biases, propagating discriminatory attributes and structural imbalances into …
Representation LearningContrastive Learning