paper-with-me

Papers

Factorised Speaker-environment Adaptive Training of Conformer Speech Recognition Systems

2023-06-26 · Jiajun Deng, Guinan Li, Xurong Xie, Zengrui Jin, Mingyu Cui, Tianzi Wang, Shujie Hu, Mengzhe Geng, Xunying Liu

Rich sources of variability in natural speech present significant challenges to current data intensive speech recognition technologies. To model both speaker and environment level diversity, this paper proposes a novel Bayesian factorised speaker-environment adaptive training and test time adaptation approach for Conformer ASR models. Speaker and environment level characteristics are separately modeled using compact hidden output transforms, which are then linearly or hierarchically combined to represent any speaker-environment combination. Bayesian learning is further utilized to model the adaptation parameter uncertainty. Experiments on the 300-hr WHAM noise corrupted Switchboard data suggest that factorised adaptation consistently outperforms the baseline and speaker label only adapted Conformers by up to 3.1% absolute (10.4% relative) word error rate reductions. Further analysis shows the proposed method offers potential for rapid adaption to unseen speaker-environment conditions.

📄 PDF Abstract BibTeX arXiv:2306.14608

Code (0)

등록된 구현이 없습니다.

Tasks

Diversityspeech-recognitionSpeech RecognitionTest-time Adaptation

Similar Papers 제목 키워드 기반

Confidence Score Based Conformer Speaker Adaptation for Speech Recognition

2022-06-24 · Jiajun Deng, Xurong Xie, Tianzi Wang, Mingyu Cui 외

A key challenge for automatic speech recognition (ASR) systems is to model the speaker level variability. In this paper, compact speaker dependent learning hidden unit contributions (LHUC) are used to facilitate both spe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

Improving the Training Recipe for a Robust Conformer-based Hybrid Model

2022-06-26 · Mohammad Zeineldeen, Jingjing Xu, Christoph Lüscher, Ralf Schlüter 외

Speaker adaptation is important to build robust automatic speech recognition (ASR) systems. In this work, we investigate various methods for speaker adaptive training (SAT) based on feature-space approaches for a conform…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

FAConformer: Frequency-Aware Convolutional Transformer for Auditory Attention Decoding

2026-06-12 · Ziwei Wang, Xingyi He, Tianwang Jia, Hongbin Wang 외 arxiv

Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a key problem for neuro-steered hearing systems. Although recent studies have achie…

Speaker-conditioning Single-channel Target Speaker Extraction using Conformer-based Architectures

2022-05-27 · Ragini Sinha, Marvin Tammen, Christian Rollwage, Simon Doclo

Target speaker extraction aims at extracting the target speaker from a mixture of multiple speakers exploiting auxiliary information about the target speaker. In this paper, we consider a complete time-domain target spea…

Target Speaker Extraction

Confidence Score Based Speaker Adaptation of Conformer Speech Recognition Systems

2023-02-15 · Jiajun Deng, Xurong Xie, Tianzi Wang, Mingyu Cui 외

Speaker adaptation techniques provide a powerful solution to customise automatic speech recognition (ASR) systems for individual users. Practical application of unsupervised model-based speaker adaptation techniques to d…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModellingSensitivity+2