paper-with-me

Papers

Xi-Vector Embedding for Speaker Recognition

2021-08-12 · Kong Aik Lee, Qiongqiong Wang, Takafumi Koshinaka

We present a Bayesian formulation for deep speaker embedding, wherein the xi-vector is the Bayesian counterpart of the x-vector, taking into account the uncertainty estimate. On the technology front, we offer a simple and straightforward extension to the now widely used x-vector. It consists of an auxiliary neural net predicting the frame-wise uncertainty of the input sequence. We show that the proposed extension leads to substantial improvement across all operating points, with a significant reduction in error rates and detection cost. On the theoretical front, our proposal integrates the Bayesian formulation of linear Gaussian model to speaker-embedding neural networks via the pooling layer. In one sense, our proposal integrates the Bayesian formulation of the i-vector to that of the x-vector. Hence, we refer to the embedding as the xi-vector, which is pronounced as /zai/ vector. Experimental results on the SITW evaluation set show a consistent improvement of over 17.5% in equal-error-rate and 10.9% in minimum detection cost.

📄 PDF Abstract BibTeX arXiv:2108.05679

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

U-vectors: Generating clusterable speaker embedding from unlabeled data

2021-02-07 · M. F. Mridha, Abu Quwsar Ohi, Muhammad Mostafa Monowar, Md. Abdul Hamid 외

Speaker recognition deals with recognizing speakers by their speech. Most speaker recognition systems are built upon two stages, the first stage extracts low dimensional correlation embeddings from speech, and the second…

Domain AdaptationSpeaker Recognition

An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS

2025-06-25 · Marie Kunešová, Zdeněk Hanzlíček, Jindřich Matoušek

Zero-shot multi-speaker text-to-speech (TTS) systems rely on speaker embeddings to synthesize speech in the voice of an unseen speaker, using only a short reference utterance. While many speaker embeddings have been deve…

Speaker Recognitiontext-to-speechText to SpeechZero-Shot Multi-Speaker TTS

Ordered and Binary Speaker Embedding

2023-05-25 · Jiaying Wang, Xianglong Wang, Namin Wang, Lantian Li 외

Modern speaker recognition systems represent utterances by embedding vectors. Conventional embedding vectors are dense and non-structural. In this paper, we propose an ordered binary embedding approach that sorts the dim…

ClusteringRetrievalSpeaker IdentificationSpeaker Recognition

Compact Speaker Embedding: lrx-vector

2020-08-11 · Munir Georges, Jonathan Huang, Tobias Bocklet

Deep neural networks (DNN) have recently been widely used in speaker recognition systems, achieving state-of-the-art performance on various benchmarks. The x-vector architecture is especially popular in this research com…

Knowledge DistillationSpeaker Recognition

Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks

2023-05-29 · Orchid Chetia Phukan, Arun Balaji Buduru, Rajesh Sharma

Speech emotion recognition (SER) is a field that has drawn a lot of attention due to its applications in diverse fields. A current trend in methods used for SER is to leverage embeddings from pre-trained models (PTMs) as…

Emotion RecognitionSpeaker RecognitionSpeech Emotion Recognition