paper-with-me

Papers

Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information

2021-10-18 · Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck

This paper contains a post-challenge performance analysis on cross-lingual speaker verification of the IDLab submission to the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). We show that current speaker embedding extractors consistently underestimate speaker similarity in within-speaker cross-lingual trials. Consequently, the typical training and scoring protocols do not put enough emphasis on the compensation of intra-speaker language variability. We propose two techniques to increase cross-lingual speaker verification robustness. First, we enhance our previously proposed Large-Margin Fine-Tuning (LM-FT) training stage with a mini-batch sampling strategy which increases the amount of intra-speaker cross-lingual samples within the mini-batch. Second, we incorporate language information in the logistic regression calibration stage. We integrate quality metrics based on soft and hard decisions of a VoxLingua107 language identification model. The proposed techniques result in a 11.7% relative improvement over the baseline model on the VoxSRC-21 test set and contributed to our third place finish in the corresponding challenge.

📄 PDF Abstract BibTeX arXiv:2110.09150

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSpeaker RecognitionSpeaker Verification

Methods 이 논문이 사용한 방법론

Test 설명 없음
Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech

2022-06-24 · Hyunjae Cho, Wonbin Jung, Junhyeok Lee, Sang Hoon Woo

In this paper, we present SANE-TTS, a stable and natural end-to-end multilingual TTS model. By the difficulty of obtaining multilingual corpus for given speaker, training multilingual TTS model with monolingual corpora i…

Rhythmtext-to-speechText to Speech

On the influence of language similarity in non-target speaker verification trials

2025-06-03 · Paul M. Reuter, Michael Jessen

In this paper, we investigate the influence of language similarity in cross-lingual non-target speaker verification trials using a state-of-the-art speaker verification system, ECAPA-TDNN, trained on multilingual and mon…

Speaker Verification

A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge

2024-06-22 · Xiaopeng Wang, Yi Lu, Xin Qi, Zhiyong Wang 외

This paper presents the development of a speech synthesis system for the LIMMITS'24 Challenge, focusing primarily on Track 2. The objective of the challenge is to establish a multi-speaker, multi-lingual Indic Text-to-Sp…

Speech Synthesistext-to-speechText to SpeechVoice Cloning

METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

2023-07-29 · Xinfa Zhu, Yi Lei, Tao Li, Yongmao Zhang 외

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotio…

DisentanglementDiversityQuantizationSpeech Synthesis+2

Combining speakers of multiple languages to improve quality of neural voices

2021-08-17 · Javier Latorre, Charlotte Bailleul, Tuuli Morrill, Alistair Conkie 외

In this work, we explore multiple architectures and training procedures for developing a multi-speaker and multi-lingual neural TTS system with the goals of a) improving the quality when the available data in the target …