paper-with-me

Papers

Time-domain speech super-resolution with GAN based modeling for telephony speaker verification

2022-09-04 · Saurabh Kataria, Jesús Villalba, Laureano Moro-Velázquez, Piotr Żelasko, Najim Dehak

Automatic Speaker Verification (ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., pooling training data from both domains, is a preferred choice for developing a universal model that works for both narrowband and wideband domains. We propose complementing this technique by performing neural upsampling of narrowband signals, also known as bandwidth extension. Our main goal is to discover and analyze high-performing time-domain Generative Adversarial Network (GAN) based models to improve our downstream state-of-the-art ASV system. We choose GANs since they (1) are powerful for learning conditional distribution and (2) allow flexible plug-in usage as a pre-processor during the training of downstream task (ASV) with data augmentation. Prior works mainly focus on feature-domain bandwidth extension and limited experimental setups. We address these limitations by 1) using time-domain extension models, 2) reporting results on three real test sets, 2) extending training data, and 3) devising new test-time schemes. We compare supervised (conditional GAN) and unsupervised GANs (CycleGAN) and demonstrate average relative improvement in Equal Error Rate of 8.6% and 7.7%, respectively. For further analysis, we study changes in spectrogram visual quality, audio perceptual quality, t-SNE embeddings, and ASV score distributions. We show that our bandwidth extension leads to phenomena such as a shift of telephone (test) embeddings towards wideband (train) signals, a negative correlation of perceptual quality with downstream performance, and condition-independent score calibration.

📄 PDF Abstract BibTeX arXiv:2209.01702

Code (0)

등록된 구현이 없습니다.

Tasks

Bandwidth ExtensionData AugmentationGenerative Adversarial NetworkSpeaker VerificationSuper-Resolution

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution

2024-09-14 · Yongjoon Lee, Chanwoo Kim

Speech Super-Resolution (SSR) is a task of enhancing low-resolution speech signals by restoring missing high-frequency components. Conventional approaches typically reconstruct log-mel features, followed by a vocoder tha…

GPUMambaSuper-Resolution

Dual-Path Cross-Modal Attention for better Audio-Visual Speech Extraction

2022-07-09 · Zhongweiyang Xu, Xulin Fan, Mark Hasegawa-Johnson

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and…

Speech ExtractionSpeech Separation

AERO: Audio Super Resolution in the Spectral Domain

2022-11-22 · Moshe Mandel, Or Tal, Yossi Adi

We present AERO, a audio super-resolution model that processes speech and music signals in the spectral domain. AERO is based on an encoder-decoder architecture with U-Net like skip connections. We optimize the model usi…

Audio Super-ResolutionBandwidth ExtensionDecoderSuper-Resolution

CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement

2022-09-22 · Sherif Abdulatif, Ruizhe Cao, Bin Yang

In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a m…

Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoder+8

A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction

2022-03-31 · Zexu Pan, Meng Ge, Haizhou Li

The speaker extraction algorithm extracts the target speech from a mixture speech containing interference speech and background noise. The extraction process sometimes over-suppresses the extracted target speech, which n…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1