paper-with-me

홈 › Papers

A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR

2024-09-09 · Giovanni Morrone, Enrico Zovato, Fabio Brugnara, Enrico Sartori, Leonardo Badino

We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a configuration file. Such flexibility allows our system to work properly in various conditions (e.g., multiple registered speakers' sets, acoustic conditions and languages) and across application domains (e.g. media monitoring, institutional, speech analytics). In this demonstration we show a practical use-case in which speaker-related information is used jointly with automatic speech recognition engines to generate speaker-attributed transcriptions. To achieve that, we employ a user-friendly web-based interface to process audio and video inputs with the chosen configuration.

📄 PDF Abstract BibTeX arXiv:2409.05750

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeaker-diarizationSpeaker DiarizationSpeaker Identificationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation

pyannote.audio: neural building blocks for speaker diarization

2019-11-04 · Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly 외

We introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-end neural building blocks that can be com…

Action DetectionActivity DetectionBIG-bench Machine LearningChange Detection+2

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

2024-03-29 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng 외

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit ade…

Self-Supervised Learningspeaker-diarizationSpeaker DiarizationSpeaker Recognition+1

A Real-time Speaker Diarization System Based on Spatial Spectrum

2021-07-20 · Siqi Zheng, Weilong Huang, Xianliang Wang, Hongbin Suo 외

In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-stan…

speaker-diarizationSpeaker DiarizationSpeaker Identification

Pretraining Multi-Speaker Identification for Neural Speaker Diarization

2025-05-30 · Shota Horiguchi, Atsushi Ando, Marc Delcroix, Naohiro Tawara

End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requiring a large amount of labeled conversa…

speaker-diarizationSpeaker DiarizationSpeaker IdentificationSpeaker Recognition