ADA-VAD: Unpaired Adversarial Domain Adaptation for Noise-Robust Voice Activity Detection
Voice Activity Detection (VAD) is becoming an essential front-end component in various speech processing systems. As those systems are commonly deployed in environments with diverse noise types and low signal-to-noise ratios (SNRs), an effective VAD method should perform robust detection of speech region out of noisy background signals. In this paper, we propose adversarial domain adaptive VAD (ADA-VAD), which is a deep neural network (DNN) based VAD method highly robust to audio samples with various noise types and low SNRs. The proposed method trains DNN models for a VAD task in a supervised manner. Simultaneously, to mitigate the performance degradation due to back-ground noises, the adversarial domain adaptation method is adopted to match the domain discrepancy between noisy and clean audio stream in an unsupervised manner. The results show that ADA-VAD achieves an average of 3.6%p and 7%p higher AUC than models trained with manually extracted features on the AVA-speech dataset and a speech database synthesized with an unseen noise database, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionDomain AdaptationSimilar Papers 제목 키워드 기반
Unsupervised Noise adaptation using Data Simulation
Deep neural network based speech enhancement approaches aim to learn a noisy-to-clean transformation using a supervised learning paradigm. However, such a trained-well transformation is vulnerable to unseen noises that a…
Domain AdaptationGenerative Adversarial NetworkSpeech EnhancementISCL: Interdependent Self-Cooperative Learning for Unpaired Image Denoising
With the advent of advances in self-supervised learning, paired clean-noisy data are no longer required in deep learning-based image denoising. However, existing blind denoising methods still require the assumption with …
DenoisingImage DenoisingSelf-Supervised LearningAn Example for Domain Adaptation Using CycleGAN
Cycle-Consistent Adversarial Network (CycleGAN) is very promising in domain adaptation. In this report, an example in medical domain will be explained. We present struecture of a CycleGAN model for unpaired image-to-imag…
Image-to-Image TranslationDomain AdaptationLearning to Clean: A GAN Perspective
In the big data era, the impetus to digitize the vast reservoirs of data trapped in unstructured scanned documents such as invoices, bank documents and courier receipts has gained fresh momentum. The scanning process oft…
DenoisingImage-to-Image TranslationTranslationNon-Local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation
Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each …
3D Human Pose EstimationDecoderPose EstimationRelation+2