SVVAD: Personal Voice Activity Detection for Speaker Verification
Voice activity detection (VAD) improves the performance of speaker verification (SV) by preserving speech segments and attenuating the effects of non-speech. However, this scheme is not ideal: (1) it fails in noisy environments or multi-speaker conversations; (2) it is trained based on inaccurate non-SV sensitive labels. To address this, we propose a speaker verification-based voice activity detection (SVVAD) framework that can adapt the speech features according to which are most informative for SV. To achieve this, we introduce a label-free training method with triplet-like losses that completely avoids the performance degradation of SV due to incorrect labeling. Extensive experiments show that SVVAD significantly outperforms the baseline in terms of equal error rate (EER) under conditions where other speakers are mixed at different ratios. Moreover, the decision boundaries reveal the importance of the different parts of speech, which are largely consistent with human judgments.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity DetectionSpeaker VerificationTripletSimilar Papers 제목 키워드 기반
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
Personalized Voice Activity Detection (PVAD) systems activate only in response to a specific target speaker. Speaker-conditioning methods are employed to inject information about the target speaker into a VAD pipeline, t…
Activity DetectionUniversal Speaker Embedding Free Target Speaker Extraction and Personal Voice Activity Detection
Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction…
Action DetectionActivity DetectionAutomatic Speech RecognitionMulti-Task Learning+5Personal VAD: Speaker-Conditioned Voice Activity Detection
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such…
Action DetectionActivity DetectionSpeaker RecognitionSpeaker Verification+2Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition
Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this …
Action DetectionActivity DetectionCPUspeech-recognition+1Array Configuration-Agnostic Personal Voice Activity Detection Based on Spatial Coherence
Personal voice activity detection has received increased attention due to the growing popularity of personal mobile devices and smart speakers. PVAD is often an integral element to speech enhancement and recognition for …
Action DetectionActivity DetectionSpeech Enhancement