paper-with-me

홈 › Papers

Learning to adapt: a meta-learning approach for speaker adaptation

2018-08-30 · Ondřej Klejch, Joachim Fainberg, Peter Bell

The performance of automatic speech recognition systems can be improved by adapting an acoustic model to compensate for the mismatch between training and testing conditions, for example by adapting to unseen speakers. The success of speaker adaptation methods relies on selecting weights that are suitable for adaptation and using good adaptation schedules to update these weights in order not to overfit to the adaptation data. In this paper we investigate a principled way of adapting all the weights of the acoustic model using a meta-learning. We show that the meta-learner can learn to perform supervised and unsupervised speaker adaptation and that it outperforms a strong baseline adapting LHUC parameters when adapting a DNN AM with 1.5M parameters. We also report initial experiments on adapting TDNN AMs, where the meta-learner achieves comparable performance with LHUC.

📄 PDF Abstract BibTeX arXiv:1808.10239

Code (1)

choko/learning_to_adapt 공식 구현 tf

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Meta-Learningspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

Meta-TTS: Meta-Learning for Few-Shot Speaker Adaptive Text-to-Speech

2021-11-07 · Sung-Feng Huang, Chyi-Jiunn Lin, Da-Rong Liu, Yi-Chen Chen 외

Personalizing a speech synthesis system is a highly desired application, where the system can generate speech with the user's voice with rare enrolled recordings. There are two main approaches to build such a system in r…

Meta-LearningSpeech Synthesistext-to-speechText to Speech

Speaker Adaptive Training using Model Agnostic Meta-Learning

2019-10-23 · Ondřej Klejch, Joachim Fainberg, Peter Bell, Steve Renals

Speaker adaptive training (SAT) of neural network acoustic models learns models in a way that makes them more suitable for adaptation to test conditions. Conventionally, model-based speaker adaptive training is performed…

Meta-Learningmodel

Meta-Voice: Fast few-shot style transfer for expressive voice cloning using meta learning

2021-11-14 · Songxiang Liu, Dan Su, Dong Yu

The task of few-shot style transfer for voice cloning in text-to-speech (TTS) synthesis aims at transferring speaking styles of an arbitrary source speaker to a target speaker's voice using very limited amount of neutral…

DisentanglementMeta-LearningStyle Transfertext-to-speech+2

OSSEM: one-shot speaker adaptive speech enhancement using meta learning

2021-11-10 · Cheng Yu, Szu-Wei Fu, Tsun-An Hsieh, Yu Tsao 외

Although deep learning (DL) has achieved notable progress in speech enhancement (SE), further research is still required for a DL-based SE system to adapt effectively and efficiently to particular speakers. In this study…

Meta-LearningSpeech Enhancement

Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation

2021-06-06 · Dongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju Hwang

With rapid progress in neural text-to-speech (TTS) models, personalized speech generation is now in high demand for many applications. For practical applicability, a TTS model should generate high-quality speech with onl…

text-to-speechText to Speech