Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding extractor acts as a weakly supervised internal VAD model and performs equally or better than comparable supervised VAD systems. Subsequently, speaker diarization can be performed efficiently by extracting the VAD logits and corresponding speaker embedding simultaneously, alleviating the need and computational overhead of an external VAD model. We provide an extensive analysis of the behavior of the frame-level attention system in current speaker verification models and propose a novel speaker diarization pipeline using ECAPA2 speaker embeddings for both VAD and embedding extraction. The proposed strategy gains state-of-the-art performance on the AMI, VoxConverse and DIHARD III diarization benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
Action DetectionActivity Detectionspeaker-diarizationSpeaker DiarizationSpeaker VerificationSimilar Papers 제목 키워드 기반
Cross modal video representations for weakly supervised active speaker localization
An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, …
Action DetectionActive Speaker LocalizationActivity DetectionEvent DetectionOnline Target Speaker Voice Activity Detection for Speaker Diarization
This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target spea…
Action DetectionActivity DetectionClusteringspeaker-diarization+1End-to-end Online Speaker Diarization with Target Speaker Tracking
This paper proposes an online target speaker voice activity detection system for speaker diarization tasks, which does not require a priori knowledge from the clustering-based diarization system to obtain the target spea…
Action DetectionActivity DetectionClusteringspeaker-diarization+1The Newsbridge -Telecom SudParis VoxCeleb Speaker Recognition Challenge 2022 System Description
We describe the system used by our team for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC 2022) in the speaker diarization track. Our solution was designed around a new combination of voice activity detection a…
Action DetectionActivity DetectionClusteringspeaker-diarization+2TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) di…
Action DetectionActivity DetectionSpeech Recognition