paper-with-me

Papers

Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information

2021-11-28 · Zhihao Du, Shiliang Zhang, Siqi Zheng, Weilong Huang, Ming Lei

Overlapping speech diarization is always treated as a multi-label classification problem. In this paper, we reformulate this task as a single-label prediction problem by encoding the multi-speaker labels with power set. Specifically, we propose the speaker embedding-aware neural diarization (SEND) method, which predicts the power set encoded labels according to the similarities between speech features and given speaker embeddings. Our method is further extended and integrated with downstream tasks by utilizing the textual information, which has not been well studied in previous literature. The experimental results show that our method achieves lower diarization error rate than the target-speaker voice activity detection. When textual information is involved, the diarization errors can be further reduced. For the real meeting scenario, our method can achieve 34.11% relative improvement compared with the Bayesian hidden Markov model based clustering algorithm.

📄 PDF Abstract BibTeX arXiv:2111.13694

Code (2)

alibaba-damo-academy/FunASR/tree/main/funasr 공식 구현 pytorch
alibaba-damo-academy/FunASR pytorch

Tasks

Action DetectionActivity DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONSpeaker Diarization

Similar Papers 제목 키워드 기반

End-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors

2020-05-20 · Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue 외

End-to-end speaker diarization for an unknown number of speakers is addressed in this paper. Recently proposed end-to-end speaker diarization outperformed conventional clustering-based speaker diarization, but it has one…

ClusteringDecoderspeaker-diarizationSpeaker Diarization

Encoder-Decoder Based Attractors for End-to-End Neural Diarization

2021-06-20 · Shota Horiguchi, Yusuke Fujita, Shinji Watanabe, Yawen Xue 외

This paper investigates an end-to-end neural diarization (EEND) method for an unknown number of speakers. In contrast to the conventional cascaded approach to speaker diarization, EEND methods are better in terms of spea…

Decoderspeaker-diarizationSpeaker Diarization

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation

EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings

2023-12-11 · Sung Hwan Mun, Min Hyun Han, Canyeong Moon, Nam Soo Kim

In recent years, there have been studies to further improve the end-to-end neural speaker diarization (EEND) systems. This letter proposes the EEND-DEMUX model, a novel framework utilizing demultiplexed speaker embedding…

speaker-diarizationSpeaker Diarization

Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors

2022-06-06 · Shota Horiguchi, Shinji Watanabe, Paola Garcia, Yuki Takashima 외

A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulatin…

Multi-Label ClassificationMUlTI-LABEL-ClASSIFICATIONspeaker-diarizationSpeaker Diarization