paper-with-me

Papers

NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization

2023-09-22 · Naohiro Tawara, Marc Delcroix, Atsushi Ando, Atsunori Ogawa

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)-based dereverberation as a front end, then applies end-to-end neural diarization with vector clustering (EEND-VC) to each channel separately. It integrates the diarization result obtained from each channel using diarization output voting error reduction plus overlap (DOVER-LAP). To harness the knowledge from the target domain and results integrated across all channels, we apply self-supervised adaptation for each session by retraining the EEND-VC with pseudo-labels derived from DOVER-LAP. The proposed system was incorporated into NTT's submission for the distant automatic speech recognition task in the CHiME-7 challenge. Our system achieved 65 % and 62 % relative improvements on development and eval sets compared to the organizer-provided VC-based baseline diarization system, securing third place in diarization performance.

📄 PDF Abstract BibTeX arXiv:2309.12656

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge

2023-08-28 · Ruoyu Wang, Maokui He, Jun Du, Hengshun Zhou 외

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios. Additionally, it also evaluates the ef…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge

2024-09-09 · Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano 외

We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-end diarization with vector clustering (EE…

Action DetectionActivity DetectionAutomatic Speech Recognitionspeech-recognition+1

Semi-supervised multi-channel speaker diarization with cross-channel attention

2023-07-17 · Shilong Wu, Jun Du, Maokui He, Shutong Niu 외

Most neural speaker diarization systems rely on sufficient manual training data labels, which are hard to collect under real-world scenarios. This paper proposes a semi-supervised speaker diarization system to utilize la…

speaker-diarizationSpeaker Diarization

Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture

2023-09-17 · Gaobin Yang, Maokui He, Shutong Niu, Ruoyu Wang 외

We propose a novel neural speaker diarization system using memory-aware multi-speaker embedding with sequence-to-sequence architecture (NSD-MS2S), which integrates the strengths of memory-aware multi-speaker embedding (M…

speaker-diarizationSpeaker Diarization

CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings

2020-04-20 · Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent 외

Following the success of the 1st, 2nd, 3rd, 4th and 5th CHiME challenges we organize the 6th CHiME Speech Separation and Recognition Challenge (CHiME-6). The new challenge revisits the previous CHiME-5 challenge and furt…

speaker-diarizationSpeaker DiarizationSpeech Enhancementspeech-recognition+2