paper-with-me

Papers

Speaker Diarization with Region Proposal Network

2020-02-14 · Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola Garcia, Yiwen Shao, Daniel Povey, Sanjeev Khudanpur

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline.

📄 PDF Abstract BibTeX arXiv:2002.06220

Code (1)

HuangZiliAndy/RPNSD 공식 구현 pytorch

Tasks

Region Proposalspeaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker

2021-08-07 · Maokui He, Desh Raj, Zili Huang, Jun Du 외

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fixed (and known) number of speakers, whic…

Action DetectionActivity DetectionRegion Proposalspeaker-diarization+1

Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios

2024-01-08 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geo…

Clustering

On the calibration of powerset speaker diarization models

2024-09-24 · Alexis Plaquet, Hervé Bredin

End-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of…

speaker-diarizationSpeaker Diarization

End-to-End Speaker Diarization as Post-Processing

2020-12-18 · Shota Horiguchi, Paola Garcia, Yusuke Fujita, Shinji Watanabe 외

This paper investigates the utilization of an end-to-end diarization model as post-processing of conventional clustering-based diarization. Clustering-based diarization methods partition frames into clusters of the numbe…

ClusteringMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONspeaker-diarization+1

MK-SGC-SC: Multiple Kernel Guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization

2026-01-24 · Nikhil Raghav, Avisek Gupta, Swagatam Das, Md Sahidullah arxiv

Speaker diarization aims to segment audio recordings into regions corresponding to individual speakers. Although unsupervised speaker diarization is inherently challenging, the prospect of identifying speaker regions wit…

Speaker Diarization