paper-with-me

Papers

Nonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs

2022-10-28 · Reo Yoneyama, Ryuichi Yamamoto, Kentaro Tachibana

Neural audio super-resolution models are typically trained on low- and high-resolution audio signal pairs. Although these methods achieve highly accurate super-resolution if the acoustic characteristics of the input data are similar to those of the training data, challenges remain: the models suffer from quality degradation for out-of-domain data, and paired data are required for training. To address these problems, we propose Dual-CycleGAN, a high-quality audio super-resolution method that can utilize unpaired data based on two connected cycle consistent generative adversarial networks (CycleGAN). Our method decomposes the super-resolution method into domain adaptation and resampling processes to handle acoustic mismatch in the unpaired low- and high-resolution signals. The two processes are then jointly optimized within the CycleGAN framework. Experimental results verify that the proposed method significantly outperforms conventional methods when paired data are not available. Code and audio samples are available from https://chomeyama.github.io/DualCycleGAN-Demo/.

📄 PDF Abstract BibTeX arXiv:2210.15887

Code (1)

chomeyama/DualCycleGAN 공식 구현 pytorch

Tasks

Audio Super-ResolutionDomain AdaptationSuper-Resolution

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Batch Normalization 설명 없음
Residual Connection 설명 없음
Instance Normalization Instance Normalization (also known as contrast normalization) is a normalization layer where: $$ y_{tijk} = \frac{x_{tijk} - \mu_{ti}}{\sqrt{\sigma_{ti}^2 +…
PatchGAN 설명 없음
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…
Tanh Activation 설명 없음
HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…

Similar Papers 제목 키워드 기반

High-quality nonparallel voice conversion based on cycle-consistent adversarial network

2018-04-02 · Fuming Fang, Junichi Yamagishi, Isao Echizen, Jaime Lorenzo-Trueba

Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data. In this paper, we…

Generative Adversarial NetworkImage-to-Image TranslationSpeech SynthesisTranslation+2

FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching

2025-01-09 · Jun-Hak Yun, Seung-bin Kim, Seong-Whan Lee

Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-b…

Audio Super-ResolutionComputational EfficiencySpeech EnhancementSuper-Resolution

AudioSR: Versatile Audio Super-resolution at Scale

2023-09-13 · Haohe Liu, Ke Chen, Qiao Tian, Wenwu Wang 외

Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope …

Audio Super-ResolutionSuper-Resolution

Audio Super-Resolution with Latent Bridge Models

2025-09-22 · Chang Li, Zehua Chen, Liyuan Wang, Jun Zhu arxiv

Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, while previous methods often suffer from su…

Audio Super-Resolution

InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation

2025-02-28 · Chong Zhang, Yukun Ma, Qian Chen, Wen Wang 외

We introduce InspireMusic, a framework integrated super resolution and large language model for high-fidelity long-form music generation. A unified framework generates high-fidelity music, songs, and audio, which incorpo…

Audio GenerationFormLanguage ModelingLanguage Modelling+3