paper-with-me

홈 › Papers

Weakly Supervised Training of Speaker Identification Models

2018-06-22 · Martin Karu, Tanel Alumäe

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is provided only at the recording level. The method uses speaker diarization to find unique speakers in each recording, and i-vectors to project the speech of each speaker to a fixed-dimensional vector. A neural network is then trained to map i-vectors to speakers, using a special objective function that allows to optimize the model using recording-level speaker labels. We report experiments on two different real-world datasets. On the VoxCeleb dataset, the method provides 94.6% accuracy on a closed set speaker identification task, surpassing the baseline performance by a large margin. On an Estonian broadcast news dataset, the method provides 66% time-weighted speaker identification recall at 93% precision.

📄 PDF Abstract BibTeX arXiv:1806.08621

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker DiarizationSpeaker Identification

Similar Papers 제목 키워드 기반

Weakly Supervised Training of Hierarchical Attention Networks for Speaker Identification

2020-05-15 · Yanpei Shi, Qiang Huang, Thomas Hain

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. In this paper, a hierarchical attention network is proposed to solve a weakly labelled speaker identification…

Speaker Identification

Advanced Rich Transcription System for Estonian Speech

2019-01-11 · Tanel Alumäe, Ottokar Tilk, Asadullah

This paper describes the current TT\"U speech transcription system for Estonian speech. The system is designed to handle semi-spontaneous speech, such as broadcast conversations, lecture recordings and interviews recorde…

Speaker Identification

Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition

2023-05-25 · Wangyou Zhang, Yanmin Qian

Self-supervised learning (SSL) based speech pre-training has attracted much attention for its capability of extracting rich representations learned from massive unlabeled data. On the other hand, the use of weakly-superv…

DenoisingSelf-Supervised Learningspeech-recognitionSpeech Recognition

Speaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker Extraction

2022-04-15 · Zifeng Zhao, Rongzhi Gu, Dongchao Yang, Jinchuan Tian 외

Dominant researches adopt supervised training for speaker extraction, while the scarcity of ideally clean corpus and channel mismatch problem are rarely considered. To this end, we propose speaker-aware mixture of mixtur…

Domain Adaptation

Cross-Talk Reduction

2024-05-30 · Zhong-Qiu Wang, Anurag Kumar, Shinji Watanabe

While far-field multi-talker mixtures are recorded, each speaker can wear a close-talk microphone so that close-talk mixtures can be recorded at the same time. Although each close-talk mixture has a high signal-to-noise …

Speech Separation