paper-with-me

홈 › Papers

Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder

2025-08-28 · Muhammad Shakeel, Yui Sudo, Yifan Peng, Chyi-Jiunn Lin, Shinji Watanabe arxiv

This paper presents a unified multi-speaker encoder (UME), a novel architecture that jointly learns representations for speaker diarization (SD), speech separation (SS), and multi-speaker automatic speech recognition (ASR) tasks using a shared speech foundational encoder. We leverage the hidden representations from multiple layers of UME as a residual weighted-sum encoding (RWSE) to effectively use information from different semantic levels, contributing to bottom-up alignment between tasks. This joint training approach captures the inherent interdependencies among the tasks, enhancing overall performance on overlapping speech data. Our evaluations demonstrate that UME substantially improves over the single-task baselines dedicated to SD, SS, and multi-speaker ASR on LibriMix evaluation sets. Notably, for SD, UME outperforms the previous studies, achieving diarization error rates of 1.37% and 2.29% on Libri2Mix and Libri3Mix evaluation sets, respectively.

📄 PDF Abstract BibTeX arXiv:2508.20474

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker DiarizationSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation

Multi-channel Conversational Speaker Separation via Neural Diarization

2023-11-15 · Hassan Taherian, DeLiang Wang

When dealing with overlapped speech, the performance of automatic speech recognition (ASR) systems substantially degrades as they are designed for single-talker speech. To enhance ASR performance in conversational or mee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognition+1

Spatial-Temporal Activity-Informed Diarization and Separation

2024-01-30 · Yicheng Hsu, Ssuhan Chen, Mingsian R. Bai

A robust multichannel speaker diarization and separation system is proposed by exploiting the spatio-temporal activity of the speakers. The system is realized in a hybrid architecture that combines the array signal proce…

speaker-diarizationSpeaker DiarizationSpeaker Separation

Online speaker diarization of meetings guided by speech separation

2024-01-30 · Elio Gruttadauria, Mathieu Fontaine, Slim Essid

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation mode…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1

Separation Guided Speaker Diarization in Realistic Mismatched Conditions

2021-07-06 · Shu-Tong Niu, Jun Du, Lei Sun, Chin-Hui Lee

We propose a separation guided speaker diarization (SGSD) approach by fully utilizing a complementarity of speech separation and speaker clustering. Since the conventional clustering-based speaker diarization (CSD) appro…

Clusteringspeaker-diarizationSpeaker DiarizationSpeech Separation