paper-with-me

Papers

Partial AUC optimization based deep speaker embeddings with class-center learning for text-independent speaker verification

2019-11-19 · Zhongxin Bai, Xiao-Lei Zhang, Jingdong Chen

Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and identification. The verification loss functions match the pipeline of speaker verification, but their implementations are difficult. Thus, most state-of-the-art deep embedding methods use the identification loss functions with softmax output units or their variants. In this paper, we propose a verification loss function, named the maximization of partial area under the Receiver-operating-characteristic (ROC) curve (pAUC), for deep embedding based text-independent speaker verification. We also propose a class-center based training trial construction method to improve the training efficiency, which is critical for the proposed loss function to be comparable to the identification loss in performance. Experiments on the Speaker in the Wild (SITW) and NIST SRE 2016 datasets show that the proposed pAUC loss function is highly competitive with the state-of-the-art identification loss functions.

📄 PDF Abstract BibTeX arXiv:1911.08077

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker VerificationText-Independent Speaker Verification

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings

2024-07-05 · Ismail Rasim Ulgen, Carlos Busso, John H. L. Hansen, Berrak Sisman

Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, the…

Speaker RecognitionSpeech SynthesisVoice Conversion

Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification

2026-01-27 · Zhihua Fang, Liang He arxiv

Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvat…

Speaker Verification

Behavioral Analysis of Pathological Speaker Embeddings of Patients During Oncological Treatment of Oral Cancer

2023-07-10 · Jenthe Thienpondt, Caroline M. Speksnijder, Kris Demuynck

In this paper, we analyze the behavior of speaker embeddings of patients during oral cancer treatment. First, we found that pre- and post-treatment speaker embeddings differ significantly, notifying a substantial change …

Speaker Verification

SAMO: Speaker Attractor Multi-Center One-Class Learning for Voice Anti-Spoofing

2022-11-04 · Siwen Ding, You Zhang, Zhiyao Duan

Voice anti-spoofing systems are crucial auxiliaries for automatic speaker verification (ASV) systems. A major challenge is caused by unseen attacks empowered by advanced speech synthesis technologies. Our previous resear…

DiversitySpeaker VerificationSpeech SynthesisVoice Anti-spoofing

Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios

2024-01-08 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geo…

Clustering