paper-with-me

Papers

3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization

2024-03-29 · Yafeng Chen, Siqi Zheng, Hui Wang, Luyao Cheng, Tinglong Zhu, Rongjie Huang, Chong Deng, Qian Chen, Shiliang Zhang, Wen Wang, Xihao Li

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system's proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker.

📄 PDF Abstract BibTeX arXiv:2403.19971

Code (2)

alibaba-damo-academy/3D-Speaker 공식 구현 pytorch
modelscope/3D-Speaker 공식 구현 pytorch

Tasks

Self-Supervised Learningspeaker-diarizationSpeaker DiarizationSpeaker RecognitionSpeaker Verification

Similar Papers 제목 키워드 기반

Kiwano: A Cutting-Edge Open-Source Toolkit for Speaker Verification

2026-06-21 · Mickael Rouvier, Pierre Michel Bousquet arxiv

In this paper, we present Kiwano, an open-source toolkit designed to advance research and evaluation for speaker verification. Kiwano provides a lightweight yet extensible framework built on PyTorch, offering standardize…

Speaker Verification

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

2024-09-24 · Shuai Wang, Ke Zhang, Shaoxiong Lin, Junjie Li 외

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increas…

Managementspeech-recognitionSpeech RecognitionTarget Speaker Extraction

SRWToolkit: An Open Source Wizard of Oz Toolkit to Create Social Robotic Avatars

2025-09-04 · Atikkhan Faridkhan Nilgar, Kristof Van Laerhoven, Ayub Kinoti arxiv

We present SRWToolkit, an open-source Wizard of Oz toolkit designed to facilitate the rapid prototyping of social robotic avatars powered by local large language models (LLMs). Our web-based toolkit enables multimodal in…

openXBOW - Introducing the Passau Open-Source Crossmodal Bag-of-Words Toolkit

2016-05-22 · Maximilian Schmitt, Björn W. Schuller

We introduce openXBOW, an open-source toolkit for the generation of bag-of-words (BoW) representations from multimodal input. In the BoW principle, word histograms were first used as features in document classification, …

Document ClassificationEmotion RecognitionGeneral ClassificationSentiment Analysis

End2You -- The Imperial Toolkit for Multimodal Profiling by End-to-End Learning

2018-02-04 · Panagiotis Tzirakis, Stefanos Zafeiriou, Bjorn W. Schuller

We introduce End2You -- the Imperial College London toolkit for multimodal profiling by end-to-end deep learning. End2You is an open-source toolkit implemented in Python and is based on Tensorflow. It provides capabiliti…

Self-Learning