paper-with-me

Papers

Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement

2026-02-04 · Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen arxiv

Pre-trained models for automatic speech recognition (ASR) and speech enhancement (SE) have exhibited remarkable capabilities under matched noise and channel conditions. However, these models often suffer from severe performance degradation when confronted with domain shifts, particularly in the presence of unseen noise and channel distortions. In view of this, we in this paper present URSA-GAN, a unified and domain-aware generative framework specifically designed to mitigate mismatches in both noise and channel conditions. URSA-GAN leverages a dual-embedding architecture that consists of a noise encoder and a channel encoder, each pre-trained with limited in-domain data to capture domain-relevant representations. These embeddings condition a GAN-based speech generator, facilitating the synthesis of speech that is acoustically aligned with the target domain while preserving phonetic content. To enhance generalization further, we propose dynamic stochastic perturbation, a novel regularization technique that introduces controlled variability into the embeddings during generation, promoting robustness to unseen domains. Empirical results demonstrate that URSA-GAN effectively reduces character error rates in ASR and improves perceptual metrics in SE across diverse noisy and mismatched channel scenarios. Notably, evaluations on compound test conditions with both channel and noise degradations confirm the generalization ability of URSA-GAN, yielding relative improvements of 16.16% in ASR performance and 15.58% in SE metrics.

📄 PDF Abstract BibTeX arXiv:2602.04307

Code (0)

등록된 구현이 없습니다.

Tasks

Speech RecognitionSpeech Enhancement

Similar Papers 제목 키워드 기반

Domain Adaptation For Formant Estimation Using Deep Learning

2016-11-06 · Yehoshua Dissen, Joseph Keshet, Jacob Goldberger, Cynthia Clopper

In this paper we present a domain adaptation technique for formant estimation using a deep network. We first train a deep learning network on a small read speech dataset. We then freeze the parameters of the trained netw…

Deep LearningDomain Adaptation

Test-Time Adaptation for Speech Emotion Recognition

2026-01-21 · Jiaheng Dong, Hong Jia, Ting Dang arxiv

The practical utility of Speech Emotion Recognition (SER) systems is undermined by their fragility to domain shifts, such as speaker variability, the distinction between acted and naturalistic emotions, and cross-corpus …

Speech Emotion RecognitionTest-time AdaptationImage ClassificationSpeech Recognition

XTREME-S: Evaluating Cross-lingual Speech Representations

2022-03-21 · Alexis Conneau, Ankur Bapna, Yu Zhang, Min Ma 외

We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages. XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and …

Representation LearningRetrievalspeech-recognitionSpeech Recognition+4

USAD: Universal Speech and Audio Representation via Distillation

2025-06-23 · Heng-Jui Chang, Saurabhchand Bhati, James Glass, Alexander H. Liu

Self-supervised learning (SSL) has revolutionized audio representations, yet models often remain domain-specific, focusing on either speech or non-speech tasks. In this work, we present Universal Speech and Audio Distill…

Audio TaggingRepresentation LearningSelf-Supervised LearningSound Classification

Non-Parametric Domain Adaptation for End-to-End Speech Translation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The end-to-end speech translation (E2E-ST) has received increasing attention due to the potential of its less error propagation, lower latency, and fewer parameters. However, the effectiveness of neural-based approaches …

Domain AdaptationTranslationTriplet