paper-with-me

Papers

MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement

2019-05-13 · Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, Shou-De Lin

Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric scores. To overcome this issue, we propose a novel MetricGAN approach with an aim to optimize the generator with respect to one or multiple evaluation metrics. Moreover, based on MetricGAN, the metric scores of the generated data can also be arbitrarily specified by users. We tested the proposed MetricGAN on a speech enhancement task, which is particularly suitable to verify the proposed approach because there are multiple metrics measuring different aspects of speech signals. Moreover, these metrics are generally complex and could not be fully optimized by Lp or conventional adversarial losses.

📄 PDF Abstract BibTeX arXiv:1905.04874

Code (5)

JasonSWFu/MetricGAN 공식 구현
anicolson/DeepXi tf
nick-nikzad/RDL-SE tf
somvy/MetricGAN
unfinity-core/MetricGAN pytorch

Tasks

Generative Adversarial NetworkSpeech Enhancement

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

2021-04-08 · Szu-Wei Fu, Cheng Yu, Tsun-An Hsieh, Peter Plantinga 외

The discrepancy between the cost function used for training a speech enhancement model and human auditory perception usually makes the quality of enhanced speech unsatisfactory. Objective evaluation metrics which conside…

Speech Enhancement

Investigating the Impact of Speech Enhancement on Audio Deepfake Detection in Noisy Environments

2026-03-16 · Anacin, Angela, Shruti Kshirsagar, Anderson R. Avila arxiv

Logical Access (LA) attacks, also known as audio deepfake attacks, use Text-to-Speech (TTS) or Voice Conversion (VC) methods to generate spoofed speech data. This can represent a serious threat to Automatic Speaker Verif…

Audio Deepfake DetectionSpeaker VerificationSpeech EnhancementVoice Conversion

MetricGAN+/-: Increasing Robustness of Noise Reduction on Unseen Data

2022-03-23 · George Close, Thomas Hain, Stefan Goetze

Training of speech enhancement systems often does not incorporate knowledge of human perception and thus can lead to unnatural sounding results. Incorporating psychoacoustically motivated speech perception metrics as par…

Speech Enhancement

Boosting Objective Scores of a Speech Enhancement Model by MetricGAN Post-processing

2020-06-18 · Szu-Wei Fu, Chien-Feng Liao, Tsun-An Hsieh, Kuo-Hsuan Hung 외

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our study applies a modified Transformer in a …

Speech Enhancement

Asymmetric GANs for Image-to-Image Translation

2019-12-14 · Hao Tang, Nicu Sebe

Existing models for unsupervised image translation with Generative Adversarial Networks (GANs) can learn the mapping from the source domain to the target domain using a cycle-consistency loss. However, these methods alwa…

Image-to-Image TranslationTranslation