paper-with-me

Papers

Investigating Confidence Estimation Measures for Speaker Diarization

2024-06-24 · Anurag Chowdhury, Abhinav Misra, Mark C. Fuhs, Monika Woszczyna

Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, background noise, and overlapping speech. These errors propagate to, and can adversely affect, downstream systems that rely on the speaker's identity, such as speaker-adapted speech recognition. One of the ways to mitigate these errors is to provide segment-level diarization confidence scores to downstream systems. In this work, we investigate multiple methods for generating diarization confidence scores, including those derived from the original diarization system and those derived from an external model. Our experiments across multiple datasets and diarization systems demonstrate that the most competitive confidence score methods can isolate ~30% of the diarization errors within segments with the lowest ~10% of confidence scores.

📄 PDF Abstract BibTeX arXiv:2406.17124

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

On the calibration of powerset speaker diarization models

2024-09-24 · Alexis Plaquet, Hervé Bredin

End-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of…

speaker-diarizationSpeaker Diarization

Speaker Diarization with Lexical Information

2020-04-13 · Tae Jin Park, Kyu J. Han, Jing Huang, Xiaodong He 외

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3

Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization

2024-03-21 · Nikhil Raghav, Md Sahidullah

Clustering speaker embeddings is crucial in speaker diarization but hasn't received as much focus as other components. Moreover, the robustness of speaker diarization across various datasets hasn't been explored when the…

Clusteringspeaker-diarizationSpeaker Diarization

Sortformer: Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens

2024-09-10 · Taejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang 외

We propose Sortformer, a novel neural model for speaker diarization, trained with unconventional objectives compared to existing end-to-end diarization models. The permutation problem in speaker diarization has long been…

speaker-diarizationSpeaker Diarization

Target-speaker Voice Activity Detection with Improved I-Vector Estimation for Unknown Number of Speaker

2021-08-07 · Maokui He, Desh Raj, Zili Huang, Jun Du 외

Target-speaker voice activity detection (TS-VAD) has recently shown promising results for speaker diarization on highly overlapped speech. However, the original model requires a fixed (and known) number of speakers, whic…

Action DetectionActivity DetectionRegion Proposalspeaker-diarization+1