paper-with-me

홈 › Papers

Designing an Effective Metric Learning Pipeline for Speaker Diarization

2018-11-01 · Vivek Sivaraman Narayanaswamy, Jayaraman J. Thiagarajan, Huan Song, Andreas Spanias

State-of-the-art speaker diarization systems utilize knowledge from external data, in the form of a pre-trained distance metric, to effectively determine relative speaker identities to unseen data. However, much of recent focus has been on choosing the appropriate feature extractor, ranging from pre-trained $i-$vectors to representations learned via different sequence modeling architectures (e.g. 1D-CNNs, LSTMs, attention models), while adopting off-the-shelf metric learning solutions. In this paper, we argue that, regardless of the feature extractor, it is crucial to carefully design a metric learning pipeline, namely the loss function, the sampling strategy and the discrimnative margin parameter, for building robust diarization systems. Furthermore, we propose to adopt a fine-grained validation process to obtain a comprehensive evaluation of the generalization power of metric learning pipelines. To this end, we measure diarization performance across different language speakers, and variations in the number of speakers in a recording. Using empirical studies, we provide interesting insights into the effectiveness of different design choices and make recommendations.

📄 PDF Abstract BibTeX arXiv:1811.00183

Code (0)

등록된 구현이 없습니다.

Tasks

Metric Learningspeaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Triplet Network with Attention for Speaker Diarization

2018-08-04 · Huan Song, Megan Willi, Jayaraman J. Thiagarajan, Visar Berisha 외

In automatic speech processing systems, speaker diarization is a crucial front-end component to separate segments from different speakers. Inspired by the recent success of deep neural networks (DNNs) in semantic inferen…

Metric Learningspeaker-diarizationSpeaker DiarizationTriplet

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

2020-11-03 · Desh Raj, Pavel Denisov, Zhuo Chen, Hakan Erdogan 외

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, spea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

An approach to optimize inference of the DIART speaker diarization pipeline

2024-08-05 · Roman Aperdannier, Sigurd Schacht, Alexander Piazza

Speaker diarization answers the question "who spoke when" for an audio file. In some diarization scenarios, low latency is required for transcription. Speaker diarization with low latency is referred to as online speaker…

Inference OptimizationKnowledge DistillationQuantizationspeaker-diarization+1

Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation

2023-09-19 · Luyao Cheng, Siqi Zheng, Qinglin Zhang, Hui Wang 외

Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and ofte…

speaker-diarizationSpeaker DiarizationSpoken Language Understanding

Speaker Diarization with Region Proposal Network

2020-02-14 · Zili Huang, Shinji Watanabe, Yusuke Fujita, Paola Garcia 외

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in vario…

Region Proposalspeaker-diarizationSpeaker Diarization