paper-with-me

Papers

Improved Vocal Effort Transfer Vector Estimation for Vocal Effort-Robust Speaker Verification

2023-05-03 · Iván López-Espejo, Santi Prieto, Alfonso Ortega, Eduardo Lleida

Despite the maturity of modern speaker verification technology, its performance still significantly degrades when facing non-neutrally-phonated (e.g., shouted and whispered) speech. To address this issue, in this paper, we propose a new speaker embedding compensation method based on a minimum mean square error (MMSE) estimator. This method models the joint distribution of the vocal effort transfer vector and non-neutrally-phonated embedding spaces and operates in a principal component analysis domain to cope with non-neutrally-phonated speech data scarcity. Experiments are carried out using a cutting-edge speaker verification system integrating a powerful self-supervised pre-trained model for speech representation. In comparison with a state-of-the-art embedding compensation method, the proposed MMSE estimator yields superior and competitive equal error rate results when tackling shouted and whispered speech, respectively.

📄 PDF Abstract BibTeX arXiv:2305.02147

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

On the Use of a Spectral Glottal Model for the Source-filter Separation of Speech

2017-12-21

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal for…

Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions

2020-08-06 · Santi Prieto, Alfonso Ortega, Iván López-Espejo, Eduardo Lleida

The performance of speaker verification systems degrades when vocal effort conditions between enrollment and test (e.g., shouted vs. normal speech) are different. This is a potential situation in non-cooperative speaker …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+1

Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise

2022-03-20 · Tuomo Raitio, Petko Petkov, Jiangchuan Li, Muhammed Shifas 외

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral …

text-to-speechText to Speech

Improving Inference-Time Optimisation for Vocal Effects Style Transfer with a Gaussian Prior

2025-05-16 · Chin-Yun Yu, Marco A. Martínez-Ramírez, Junghyun Koo, Wei-Hsiang Liao 외

Style Transfer with Inference-Time Optimisation (ST-ITO) is a recent approach for transferring the applied effects of a reference audio to a raw audio track. It optimises the effect parameters to minimise the distance be…

Style Transfer

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

2026-06-25 · Zahra Omidi, John H. L. Hansen arxiv

The variations in vocal effort range (e.g. whisper, soft, neutral, loud, shout) alter production and speech acoustics, reducing intelligibility and limiting the robustness of any subsequent speech technology. Classificat…

Data Augmentation