A Unified Deep Speaker Embedding Framework for Mixed-Bandwidth Speech Data
This paper proposes a unified deep speaker embedding framework for modeling speech data with different sampling rates. Considering the narrowband spectrogram as a sub-image of the wideband spectrogram, we tackle the joint modeling problem of the mixed-bandwidth data in an image classification manner. From this perspective, we elaborate several mixed-bandwidth joint training strategies under different training and test data scenarios. The proposed systems are able to flexibly handle the mixed-bandwidth speech data in a single speaker embedding model without any additional downsampling, upsampling, bandwidth extension, or padding operations. We conduct extensive experimental studies on the VoxCeleb1 dataset. Furthermore, the effectiveness of the proposed approach is validated by the SITW and NIST SRE 2016 datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Bandwidth Extensionimage-classificationImage ClassificationSimilar Papers 제목 키워드 기반
Bandwidth Embeddings for Mixed-bandwidth Speech Recognition
In this paper, we tackle the problem of handling narrowband and wideband speech by building a single acoustic model (AM), also called mixed bandwidth AM. In the proposed approach, an auxiliary input feature is used to pr…
speech-recognitionSpeech RecognitionTime-domain speech super-resolution with GAN based modeling for telephony speaker verification
Automatic Speaker Verification (ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., …
Bandwidth ExtensionData AugmentationGenerative Adversarial NetworkSpeaker Verification+1Single microphone speaker extraction using unified time-frequency Siamese-Unet
In this paper we present a unified time-frequency method for speaker extraction in clean and noisy conditions. Given a mixed signal, along with a reference signal, the common approaches for extracting the desired speaker…
blind source separationDecoderBuilding a mixed-lingual neural TTS system with only monolingual data
When deploying a Chinese neural text-to-speech (TTS) synthesis system, one of the challenges is to synthesize Chinese utterances with English phrases or words embedded. This paper looks into the problem in the encoder-de…
Decodertext-to-speechText to SpeechLearning Singing From Speech
We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unifi…
Speech SynthesisVoice Conversion