paper-with-me

Papers

Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

2025-05-22 · Ming Cheng, Fei Su, Cancan Li, Juan Liu, Ming Li

This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework to generate initial predictions using single-channel audio. Then, we extend the original S2SND framework to create a new version, Multi-Channel Sequence-to-Sequence Neural Diarization (MC-S2SND), which refines the initial results using multi-channel audio. The final system achieves a diarization error rate (DER) of 8.09% on the evaluation set of the competition database, ranking first place in the speaker diarization task of the MISP 2025 Challenge.

📄 PDF Abstract BibTeX arXiv:2505.16387

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Spatial-aware Speaker Diarization for Multi-channel Multi-party Meeting

2022-09-24 · Jie Wang, Yuji Liu, Binling Wang, Yiming Zhi 외

This paper describes a spatial-aware speaker diarization system for the multi-channel multi-party meeting. The diarization system obtains direction information of speaker by microphone array. Speaker spatial embedding is…

speaker-diarizationSpeaker Diarization

The xmuspeech system for multi-channel multi-party meeting transcription challenge

2022-02-11 · Jie Wang, Yuji Liu, Binling Wang, Yiming Zhi 외

This paper describes the system developed by the XMUSPEECH team for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT). For the speaker diarization task, we propose a multi-channel speaker diarization …

speaker-diarizationSpeaker Diarization

Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors

2021-07-04 · Shota Horiguchi, Shinji Watanabe, Paola Garcia, Yawen Xue 외

Attractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the cas…

Clustering

Neural Speaker Diarization with Speaker-Wise Chain Rule

2020-06-02 · Yusuke Fujita, Shinji Watanabe, Shota Horiguchi, Yawen Xue 외

Speaker diarization is an essential step for processing multi-speaker audio. Although an end-to-end neural diarization (EEND) method achieved state-of-the-art performance, it is limited to a fixed number of speakers. In …

speaker-diarizationSpeaker Diarization

Multimodal Speaker Segmentation and Diarization using Lexical and Acoustic Cues via Sequence to Sequence Neural Networks

2018-05-28

While there has been substantial amount of work in speaker diarization recently, there are few efforts in jointly employing lexical and acoustic information for speaker segmentation. Towards that, we investigate a speake…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+2