paper-with-me

Papers

Exploring Gaussian mixture model framework for speaker adaptation of deep neural network acoustic models

2020-03-15 · Natalia Tomashenko, Yuri Khokhlov, Yannick Esteve

In this paper we investigate the GMM-derived (GMMD) features for adaptation of deep neural network (DNN) acoustic models. The adaptation of the DNN trained on GMMD features is done through the maximum a posteriori (MAP) adaptation of the auxiliary GMM model used for GMMD feature extraction. We explore fusion of the adapted GMMD features with conventional features, such as bottleneck and MFCC features, in two different neural network architectures: DNN and time-delay neural network (TDNN). We analyze and compare different types of adaptation techniques such as i-vectors and feature-space adaptation techniques based on maximum likelihood linear regression (fMLLR) with the proposed adaptation approach, and explore their complementarity using various types of fusion such as feature level, posterior level, lattice level and others in order to discover the best possible way of combination. Experimental results on the TED-LIUM corpus show that the proposed adaptation technique can be effectively integrated into DNN and TDNN setups at different levels and provide additional gain in recognition performance: up to 6% of relative word error rate reduction (WERR) over the strong feature-space adaptation techniques based on maximum likelihood linear regression (fMLLR) speaker adapted DNN baseline, and up to 18% of relative WERR in comparison with a speaker independent (SI) DNN baseline model, trained on conventional features. For TDNN models the proposed approach achieves up to 26% of relative WERR in comparison with a SI baseline, and up 13% in comparison with the model adapted by using i-vectors. The analysis of the adapted GMMD features from various points of view demonstrates their effectiveness at different levels.

📄 PDF Abstract BibTeX arXiv:2003.06894

Code (0)

등록된 구현이 없습니다.

Tasks

regression

Methods 이 논문이 사용한 방법론

Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

ADAPTERMIX: Exploring the Efficacy of Mixture of Adapters for Low-Resource TTS Adaptation

2023-05-29 · Ambuj Mehrish, Abhinav Ramesh Kashyap, Li Yingting, Navonil Majumder 외

There are significant challenges for speaker adaptation in text-to-speech for languages that are not widely spoken or for speakers with accents or dialects that are not well-represented in the training data. To address t…

Speech Synthesistext-to-speechText to Speech

Exploring Speaker Diarization with Mixture of Experts

2025-06-17 · Gaobin Yang, Maokui He, Shutong Niu, Ruoyu Wang 외

In this paper, we propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates a memory-aware multi-speaker embedding mo…

Mixture-of-Expertsspeaker-diarizationSpeaker Diarization

MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition

2025-05-30 · Chengxi Deng, Xurong Xie, Shujie Hu, Mengzhe Geng 외

This paper proposes a novel Mixture of Prompt-Experts based Speaker Adaptation approach (MOPSA) for elderly speech recognition. It allows zero-shot, real-time adaptation to unseen speakers, and leverages domain knowledge…

Decoderspeech-recognitionSpeech Recognition

Speaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker Extraction

2022-04-15 · Zifeng Zhao, Rongzhi Gu, Dongchao Yang, Jinchuan Tian 외

Dominant researches adopt supervised training for speaker extraction, while the scarcity of ideally clean corpus and channel mismatch problem are rarely considered. To this end, we propose speaker-aware mixture of mixtur…

Domain Adaptation

Speaker Identification From Youtube Obtained Data

2014-11-11 · Nitesh Kumar Chaudhary

An efficient, and intuitive algorithm is presented for the identification of speakers from a long dataset (like YouTube long discussion, Cocktail party recorded audio or video).The goal of automatic speaker identificatio…

parameter estimationQuantizationSpeaker IdentificationSpeaker Verification