paper-with-me

Papers

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

2024-09-19 · Chien-Chun Wang, Li-Wei Chen, Cheng-Kang Chou, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics.

📄 PDF Abstract BibTeX arXiv:2409.12386

Code (1)

jethrowangsir/cada-gan 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Generative Adversarial NetworkRobust Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Learning Semantic-aware Normalization for Generative Adversarial Networks

2020-12-01 · NeurIPS 2020 12 · Heliang Zheng, Jianlong Fu, Yanhong Zeng, Jiebo Luo 외

The recent advances in image generation have been achieved by style-based image generators. Such approaches learn to disentangle latent factors in different image scales and encode latent factors as “style” to control im…

Image GenerationImage InpaintingUnconditional Image Generation

DTGAN: Dual Attention Generative Adversarial Networks for Text-to-Image Generation

2020-11-05 · Zhenxing Zhang, Lambert Schomaker

Most existing text-to-image generation methods adopt a multi-stage modular architecture which has three significant problems: 1) Training multiple networks increases the run time and affects the convergence and stability…

Generative Adversarial NetworkImage GenerationSentenceText to Image Generation+1

Unpaired Image Enhancement with Quality-Attention Generative Adversarial Network

2020-12-30 · Zhangkai Ni, Wenhan Yang, Shiqi Wang, Lin Ma 외

In this work, we aim to learn an unpaired image enhancement model, which can enrich low-quality images with the characteristics of high-quality images provided by users. We propose a quality attention generative adversar…

Generative Adversarial NetworkImage Enhancement

Multisource Collaborative Domain Generalization for Cross-Scene Remote Sensing Image Classification

2024-12-05 · Zhu Han, Ce Zhang, Lianru Gao, Zhiqiang Zeng 외

Cross-scene image classification aims to transfer prior knowledge of ground materials to annotate regions with different distributions and reduce hand-crafted cost in the field of remote sensing. However, existing approa…

DiversityDomain Generalizationimage-classificationImage Classification+3

Self Sparse Generative Adversarial Networks

2021-01-26 · Wenliang Qian, Yang Xu, WangMeng Zuo, Hui Li

Generative Adversarial Networks (GANs) are an unsupervised generative model that learns data distribution through adversarial training. However, recent experiments indicated that GANs are difficult to train due to the re…

Generative Adversarial NetworkImage Generation