paper-with-me

홈 › Papers

Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks

2019-10-01 · Kin Wai Cheuk, Balamurali B. T., Gemma Roig, Dorien Herremans

We present an approach to tackle the speaker recognition problem using Triplet Neural Networks. Currently, the $i$-vector representation with probabilistic linear discriminant analysis (PLDA) is the most commonly used technique to solve this problem, due to high classification accuracy with a relatively short computation time. In this paper, we explore a neural network approach, namely Triplet Neural Networks (TNNs), to built a latent space for different classifiers to solve the Multi-Target Speaker Detection and Identification Challenge Evaluation 2018 (MCE 2018) dataset. This training set contains $i$-vectors from 3,631 speakers, with only 3 samples for each speaker, thus making speaker recognition a challenging task. When using the train and development set for training both the TNN and baseline model (i.e., similarity evaluation directly on the $i$-vector representation), our proposed model outperforms the baseline by 23%. When reducing the training data to only using the train set, our method results in 309 confusions for the Multi-target speaker identification task, which is 46% better than the baseline model. These results show that the representational power of TNNs is especially evident when training on small datasets with few instances available per class.

📄 PDF Abstract BibTeX arXiv:1910.01463

Code (1)

KinWaiCheuk/MCE2018 공식 구현 tf

Tasks

Speaker IdentificationSpeaker RecognitionTriplet

Similar Papers 제목 키워드 기반

Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations

2018-04-09 · Ju-chieh Chou, Cheng-chieh Yeh, Hung-Yi Lee, Lin-shan Lee

Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for ea…

DecoderVoice Conversion

Breaking Audio Large Language Models by Attacking Only the Encoder: A Universal Targeted Latent-Space Audio Attack

2025-12-29 · Roee Ziv, Raz Lapid, Moshe Sipper arxiv

Audio-language models combine audio encoders with large language models to enable multimodal reasoning, but they also introduce new security vulnerabilities. We propose a universal targeted latent space attack, an encode…

Multimodal ReasoningAdversarial Attack

DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning

2020-12-12 · Mufan Sang, Wei Xia, John H. L. Hansen

Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, domain mismatch is still a challenging problem in this field. In this study, we propose a novel …

DisentanglementDomain AdaptationRepresentation LearningSpeaker Verification

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline

2025-05-25 · Helin Wang, Jiarui Hai, Dongchao Yang, Chen Chen 외

Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent a…

Speech ExtractionSpeech Separation

VAE-based Domain Adaptation for Speaker Verification

2019-08-27 · Xueyi Wang, Lantian Li, Dong Wang

Deep speaker embedding has achieved satisfactory performance in speaker verification. By enforcing the neural model to discriminate the speakers in the training set, deep speaker embedding (called `x-vectors`) can be der…

Domain AdaptationSpeaker Verification