paper-with-me

홈 › Papers

Cross-Channel Attention-Based Target Speaker Voice Activity Detection: Experimental Results for M2MeT Challenge

2022-02-06 · Weiqing Wang, Xiaoyi Qin, Ming Li

In this paper, we present the speaker diarization system for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) from team DKU_DukeECE. As the highly overlapped speech exists in the dataset, we employ an x-vector-based target-speaker voice activity detection (TS-VAD) to find the overlap between speakers. For the single-channel scenario, we separately train a model for each of the 8 channels and fuse the results. We also employ the cross-channel self-attention to further improve the performance, where the non-linear spatial correlations between different channels are learned and fused. Experimental results on the evaluation set show that the single-channel TS-VAD reduces the DER by over 75% from 12.68\% to 3.14%. The multi-channel TS-VAD further reduces the DER by 28% and achieves a DER of 2.26%. Our final submitted system achieves a DER of 2.98% on the AliMeeting test set, which ranks 1st in the M2MET challenge.

📄 PDF Abstract BibTeX arXiv:2202.02687

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

2023-12-14 · Junjie Li, Yiwei Guo, Xie Chen, Kai Yu

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the voice of generated speech can be controll…

PositionVoice Conversion

Beamformer-Guided Target Speaker Extraction

2023-03-15 · Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …

Target Speaker Extraction

L-SpEx: Localized Target Speaker Extraction

2022-02-21 · Meng Ge, Chenglin Xu, Longbiao Wang, Eng Siong Chng 외

Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…

Target Speaker Extraction

Individualized Conditioning and Negative Distances for Speaker Separation

2022-10-12 · Tao Sun, Nidal Abuhajar, Shuyu Gong, Zhewei Wang 외

Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning …

Speaker SeparationTriplet

Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario

2020-05-14 · Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach, Yuri Khokhlov 외

Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainly due to the limited ability to handle o…

Action DetectionActivity DetectionBinary ClassificationClustering+2