paper-with-me

홈 › Papers

Developing Far-Field Speaker System Via Teacher-Student Learning

2018-04-14 · Jinyu Li, Rui Zhao, Zhuo Chen, Changliang Liu, Xiong Xiao, Guoli Ye, Yifan Gong

In this study, we develop the keyword spotting (KWS) and acoustic model (AM) components in a far-field speaker system. Specifically, we use teacher-student (T/S) learning to adapt a close-talk well-trained production AM to far-field by using parallel close-talk and simulated far-field data. We also use T/S learning to compress a large-size KWS model into a small-size one to fit the device computational cost. Without the need of transcription, T/S learning well utilizes untranscribed data to boost the model performance in both the AM adaptation and KWS model compression. We further optimize the models with sequence discriminative training and live data to reach the best performance of systems. The adapted AM improved from the baseline by 72.60% and 57.16% relative word error rate reduction on play-back and live test data, respectively. The final KWS model size was reduced by 27 times from a large-size KWS model without losing accuracy.

📄 PDF Abstract BibTeX arXiv:1804.05166

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword SpottingModel Compression

Methods 이 논문이 사용한 방법론

AM 설명 없음

Similar Papers 제목 키워드 기반

A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures

2023-06-01 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-student approach is employed: the teacher c…

Open-set Short Utterance Forensic Speaker Verification using Teacher-Student Network with Explicit Inductive Bias

2020-09-21 · Mufan Sang, Wei Xia, John H. L. Hansen

In forensic applications, it is very common that only small naturalistic datasets consisting of short utterances in complex or unknown acoustic environments are available. In this study, we propose a pipeline solution to…

Inductive BiasKnowledge DistillationSpeaker Verification

Short utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings

2018-10-25 · Jee-weon Jung, Hee-Soo Heo, Hye-jin Shim, Ha-Jin Yu

The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification s…

Speaker VerificationText-Independent Speaker Verification

Data-augmented cross-lingual synthesis in a teacher-student framework

2022-03-31 · Marcel de Korte, Jaebok Kim, Aki Kunikoshi, Adaeze Adigwe 외

Cross-lingual synthesis can be defined as the task of letting a speaker generate fluent synthetic speech in another language. This is a challenging task, and resulting speech can suffer from reduced naturalness, accented…

Conditional Teacher-Student Learning

2019-04-28 · Zhong Meng, Jinyu Li, Yong Zhao, Yifan Gong

The teacher-student (T/S) learning has been shown to be effective for a variety of problems such as domain adaptation and model compression. One shortcoming of the T/S learning is that a teacher model, not always perfect…

Domain AdaptationModel Compression