paper-with-me

Papers

Royalflush Speaker Diarization System for ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge

2022-02-10 · Jingguang Tian, Xinhui Hu, Xinkang Xu

This paper describes the Royalflush speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription Challenge(M2MeT). Our system comprises speech enhancement, overlapped speech detection, speaker embedding extraction, speaker clustering, speech separation and system fusion. In this system, we made three contributions. First, we propose an architecture of combining the multi-channel and U-Net-based models, aiming at utilizing the benefits of these two individual architectures, for far-field overlapped speech detection. Second, in order to use overlapped speech detection model to help speaker diarization, a speech separation based overlapped speech handling approach, in which the speaker verification technique is further applied, is proposed. Third, we explore three speaker embedding methods, and obtained the state-of-the-art performance on the CNCeleb-E test set. With these proposals, our best individual system significantly reduces DER from 15.25% to 6.40%, and the fusion of four systems finally achieves a DER of 6.30% on the far-field Alimeeting evaluation set.

📄 PDF Abstract BibTeX arXiv:2202.04814

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker DiarizationSpeaker VerificationSpeech EnhancementSpeech Separation

Similar Papers 제목 키워드 기반

The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

2022-02-04 · Naijun Zheng, Na Li, Xixin Wu, Lingwei Meng 외

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization an…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

The Volcspeech system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

2022-02-09 · Chen Shen, Yi Liu, Wenzhi Fan, Bin Wang 외

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clustering-based speaker diarization system …

Data AugmentationLanguage Modellingspeaker-diarizationSpeaker Diarization+2

The USTC-Ximalaya system for the ICASSP 2022 multi-channel multi-party meeting transcription (M2MeT) challenge

2022-02-10 · Maokui He, Xiang Lv, Weilin Zhou, JingJing Yin 외

We propose two improvements to target-speaker voice activity detection (TS-VAD), the core component in our proposed speaker diarization system that was submitted to the 2022 Multi-Channel Multi-Party Meeting Transcriptio…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization

The Royalflush System for VoxCeleb Speaker Recognition Challenge 2022

2022-09-19 · Jingguang Tian, Xinhui Hu, Xinkang Xu

In this technical report, we describe the Royalflush submissions for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our submissions contain track 1, which is for supervised speaker verification and track 3,…

ClusteringDomain AdaptationSpeaker RecognitionSpeaker Verification

The RoyalFlush System of Speech Recognition for M2MeT Challenge

2022-02-03 · Shuaishuai Ye, Peiyao Wang, Shunfei Chen, Xinhui Hu 외

This paper describes our RoyalFlush system for the track of multi-speaker automatic speech recognition (ASR) in the M2MeT challenge. We adopted the serialized output training (SOT) based multi-speakers ASR system with la…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+4