paper-with-me

Papers

Unified Hypersphere Embedding for Speaker Recognition

2018-07-22 · Mahdi Hajibabaei, Dengxin Dai

Incremental improvements in accuracy of Convolutional Neural Networks are usually achieved through use of deeper and more complex models trained on larger datasets. However, enlarging dataset and models increases the computation and storage costs and cannot be done indefinitely. In this work, we seek to improve the identification and verification accuracy of a text-independent speaker recognition system without use of extra data or deeper and more complex models by augmenting the training and testing data, finding the optimal dimensionality of embedding space and use of more discriminative loss functions. Results of experiments on VoxCeleb dataset suggest that: (i) Simple repetition and random time-reversion of utterances can reduce prediction errors by up to 18%. (ii) Lower dimensional embeddings are more suitable for verification. (iii) Use of proposed logistic margin loss function leads to unified embeddings with state-of-the-art identification and competitive verification accuracies.

📄 PDF Abstract BibTeX arXiv:1807.08312

Code (1)

MahdiHajibabaei/unified-embedding 공식 구현 caffe2

Tasks

Speaker RecognitionText-Independent Speaker Recognition

Similar Papers 제목 키워드 기반

Deep Speaker: an End-to-End Neural Speaker Embedding System

2017-05-05 · Chao Li, Xiaokong Ma, Bing Jiang, Xiangang Li 외

We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity. The embeddings generated by Deep Speaker can be used for many ta…

ClusteringSpeaker IdentificationSpeaker RecognitionTriplet

Probabilistic Spherical Discriminant Analysis: An Alternative to PLDA for length-normalized embeddings

2022-03-28 · Niko Brümmer, Albert Swart, Ladislav Mošner, Anna Silnova 외

In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA. Both have advantages and disadvantages, depending on …

Speaker Recognition

Toroidal Probabilistic Spherical Discriminant Analysis

2022-10-27 · Anna Silnova, Niko Brümmer, Albert Swart, Lukáš Burget

In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring back-ends are commonly used, namely cosine scoring and PLDA. We have recently proposed PSDA, an analog to PLDA t…

FormSpeaker Recognition

Investigation of Speaker Representation for Target-Speaker Speech Processing

2024-10-15 · Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng 외

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting infor…

Action DetectionActivity DetectionAutomatic Speech RecognitionSpeaker Recognition+4

Modeling Named Entity Embedding Distribution into Hypersphere

2019-09-03 · Zhuosheng Zhang, Bingjie Tang, Zuchao Li, Hai Zhao

This work models named entity distribution from a way of visualizing topological structure of embedding space, so that we make an assumption that most, if not all, named entities (NEs) for a language tend to aggregate to…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)