paper-with-me

Papers

PAS: Partial Additive Speech Data Augmentation Method for Noise Robust Speaker Verification

2023-07-20 · Wonbin Kim, Hyun-seo Shin, Ju-ho Kim, Jungwoo Heo, Chan-yeong Lim, Ha-Jin Yu

Background noise reduces speech intelligibility and quality, making speaker verification (SV) in noisy environments a challenging task. To improve the noise robustness of SV systems, additive noise data augmentation method has been commonly used. In this paper, we propose a new additive noise method, partial additive speech (PAS), which aims to train SV systems to be less affected by noisy environments. The experimental results demonstrate that PAS outperforms traditional additive noise in terms of equal error rates (EER), with relative improvements of 4.64% and 5.01% observed in SE-ResNet34 and ECAPA-TDNN. We also show the effectiveness of proposed method by analyzing attention modules and visualizing speaker embeddings.

📄 PDF Abstract BibTeX arXiv:2307.10628

Code (1)

rst0070/Partial_Additive_Speech 공식 구현 pytorch

Tasks

Data AugmentationSpeaker Verification

Similar Papers 제목 키워드 기반

Data augmentation using prosody and false starts to recognize non-native children's speech

2020-08-29 · Hemant Kathania, Mittul Singh, Tamás Grósz, Mikko Kurimo

This paper describes AaltoASR's speech recognition system for the INTERSPEECH 2020 shared task on Automatic Speech Recognition (ASR) for non-native children's speech. The task is to recognize non-native speech from child…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationLanguage Modeling+3

Data Augmenting Contrastive Learning of Speech Representations in the Time Domain

2020-07-02 · Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf 외

Contrastive Predictive Coding (CPC), based on predicting future segments of speech based on past segments is emerging as a powerful algorithm for representation learning of speech signal. However, it still under-performs…

Contrastive LearningData AugmentationRepresentation Learning

End-to-end Recurrent Denoising Autoencoder Embeddings for Speaker Identification

2020-03-13 · Esther Rituerto-González, Carmen Peláez-Moreno

Speech 'in-the-wild' is a handicap for speaker recognition systems due to the variability induced by real-life conditions, such as environmental noise and the emotional state of the speaker. Taking advantage of the princ…

Data AugmentationDenoisingRepresentation LearningSpeaker Identification+1

Partial differential equation regularization for supervised machine learning

2019-10-03 · Adam M. Oberman

This article is an overview of supervised machine learning problems for regression and classification. Topics include: kernel methods, training by stochastic gradient descent, deep learning architecture, losses for class…

BIG-bench Machine LearningClassificationData AugmentationDeep Learning+4

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

2026-06-25 · Zahra Omidi, John H. L. Hansen arxiv

The variations in vocal effort range (e.g. whisper, soft, neutral, loud, shout) alter production and speech acoustics, reducing intelligibility and limiting the robustness of any subsequent speech technology. Classificat…

Data Augmentation