Self-Adaptive Soft Voice Activity Detection using Deep Neural Networks for Robust Speaker Verification
Voice activity detection (VAD), which classifies frames as speech or non-speech, is an important module in many speech applications including speaker verification. In this paper, we propose a novel method, called self-adaptive soft VAD, to incorporate a deep neural network (DNN)-based VAD into a deep speaker embedding system. The proposed method is a combination of the following two approaches. The first approach is soft VAD, which performs a soft selection of frame-level features extracted from a speaker feature extractor. The frame-level features are weighted by their corresponding speech posteriors estimated from the DNN-based VAD, and then aggregated to generate a speaker embedding. The second approach is self-adaptive VAD, which fine-tunes the pre-trained VAD on the speaker verification data to reduce the domain mismatch. Here, we introduce two unsupervised domain adaptation (DA) schemes, namely speech posterior-based DA (SP-DA) and joint learning-based DA (JL-DA). Experiments on a Korean speech database demonstrate that the verification performance is improved significantly in real-world environments by using self-adaptive soft VAD.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionDomain AdaptationSpeaker VerificationUnsupervised Domain AdaptationSimilar Papers 제목 키워드 기반
An End-to-End Architecture for Keyword Spotting and Voice Activity Detection
We propose a single neural network architecture for two tasks: on-line keyword spotting and voice activity detection. We develop novel inference algorithms for an end-to-end Recurrent Neural Network trained with the Conn…
Action DetectionActivity DetectionGeneral ClassificationKeyword SpottingSelf-supervised Pretraining for Robust Personalized Voice Activity Detection in Adverse Conditions
In this paper, we propose the use of self-supervised pretraining on a large unlabelled data set to improve the performance of a personalized voice activity detection (VAD) model in adverse conditions. We pretrain a long …
Action DetectionActivity DetectionDenoisingMultitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0
Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2…
Action DetectionActivity DetectionChange DetectionSelf-Supervised LearningVoice Activity Projection: Self-supervised Learning of Turn-taking Events
The modeling of turn-taking in dialog can be viewed as the modeling of the dynamics of voice activity of the interlocutors. We extend prior work and define the predictive task of Voice Activity Projection, a general, sel…
Self-Supervised LearningCross-domain Voice Activity Detection with Self-Supervised Representations
Voice Activity Detection (VAD) aims at detecting speech segments on an audio signal, which is a necessary first step for many today's speech based applications. Current state-of-the-art methods focus on training a neural…
Action DetectionActivity DetectionSelf-Supervised Learning