Papers Audio Super-Resolution
“Audio Super-Resolution” 태그가 달린 논문 28편 · 필터 해제
FastWave: Optimized Diffusion Model for Audio Super-Resolution
Audio Super-Resolution is a set of techniques aimed at high-quality estimation of the given signal as if it would be sampled with higher sample rate. Among suggested methods there are diffusion and flow models (which are…
Audio Super-ResolutionDiscriminating real and synthetic super-resolved audio samples using embedding-based classifiers
Generative adversarial networks (GANs) and diffusion models have recently achieved state-of-the-art performance in audio super-resolution (ADSR), producing perceptually convincing wideband audio from narrowband inputs. H…
Audio Super-ResolutionHQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
Zero-shot singing voice conversion (SVC) transforms a source singer's timbre to an unseen target speaker's voice while preserving melodic content without fine-tuning. Existing methods model speaker timbre and vocal conte…
Audio Super-ResolutionVoice ConversionUniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
In this paper, we present a vocoder-free framework for audio super-resolution that employs a flow matching generative model to capture the conditional distribution of complex-valued spectral coefficients. Unlike conventi…
Audio Super-ResolutionAudio Super-Resolution with Latent Bridge Models
Audio super-resolution (SR), i.e., upsampling the low-resolution (LR) waveform to the high-resolution (HR) version, has recently been explored with diffusion and bridge models, while previous methods often suffer from su…
Audio Super-ResolutionInference-time Scaling for Diffusion-based Audio Super-resolution
Diffusion models have demonstrated remarkable success in generative tasks, including audio super-resolution (SR). In many applications like movie post-production and album mastering, substantial computational budgets are…
Audio Super-ResolutionFlashSR: One-step Versatile Audio Super-resolution via Diffusion Distillation
Versatile audio super-resolution (SR) is the challenging task of restoring high-frequency components from low-resolution audio with sampling rates between 4kHz and 32kHz in various domains such as music, speech, and soun…
Audio Super-ResolutionSuper-ResolutionFLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching
Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-b…
Audio Super-ResolutionComputational EfficiencySpeech EnhancementSuper-ResolutionAEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
Audio super-resolution aims to enhance low-resolution signals by creating high-frequency content. In this work, we modify the architecture of AERO (a state-of-the-art system for this task) for music super-resolution. SPe…
Audio Super-ResolutionGPUMambaState Space Models+1Gull: A Generative Multifunctional Audio Codec
We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time com…
Audio CompressionAudio Source SeparationAudio Super-ResolutionDecoder+2AudioSR: Versatile Audio Super-resolution at Scale
Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope …
Audio Super-ResolutionSuper-ResolutionEdge Storage Management Recipe with Zero-Shot Data Compression for Road Anomaly Detection
Recent studies show edge computing-based road anomaly detection systems which may also conduct data collection simultaneously. However, the edge computers will have small data storage but we need to store the collected a…
Anomaly DetectionAudio CompressionAudio Super-ResolutionData Compression+3AERO: Audio Super Resolution in the Spectral Domain
We present AERO, a audio super-resolution model that processes speech and music signals in the spectral domain. AERO is based on an encoder-decoder architecture with U-Net like skip connections. We optimize the model usi…
Audio Super-ResolutionBandwidth ExtensionDecoderSuper-ResolutionNonparallel High-Quality Audio Super Resolution with Domain Adaptation and Resampling CycleGANs
Neural audio super-resolution models are typically trained on low- and high-resolution audio signal pairs. Although these methods achieve highly accurate super-resolution if the acoustic characteristics of the input data…
Audio Super-ResolutionDomain AdaptationSuper-ResolutionCMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement
In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a m…
Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoder+8NU-Wave 2: A General Neural Audio Upsampling Model for Various Sampling Rates
Conventionally, audio super-resolution models fixed the initial and the target sampling rates, which necessitate the model to be trained for each pair of sampling rates. We introduce NU-Wave 2, a diffusion model for neur…
Audio Super-ResolutionSuper-ResolutionNeural Vocoder is All You Need for Speech Super-resolution
Speech super-resolution (SR) is a task to increase speech sampling rate by generating high-frequency components. Existing speech SR methods are trained in constrained experimental settings, such as a fixed upsampling rat…
AllAudio Super-ResolutionBandwidth ExtensionSuper-ResolutionLearning Continuous Representation of Audio for Arbitrary Scale Super Resolution
Audio super resolution aims to predict the missing high resolution components of the low resolution audio signals. While audio in nature is a continuous signal, current approaches treat it as discrete data (i.e., input i…
Audio Super-ResolutionSelf-Supervised LearningSuper-ResolutionTUNet: A Block-online Bandwidth Extension Model based on Transformers and Self-supervised Pretraining
We introduce a block-online variant of the temporal feature-wise linear modulation (TFiLM) model to achieve bandwidth extension. The proposed architecture simplifies the UNet backbone of the TFiLM to reduce inference tim…
Audio Super-ResolutionBandwidth ExtensionSensitivityAn investigation of pre-upsampling generative modelling and Generative Adversarial Networks in audio super resolution
There have been several successful deep learning models that perform audio super-resolution. Many of these approaches involve using preprocessed feature extraction which requires a lot of domain-specific signal processin…
Audio GenerationAudio Super-ResolutionImage Super-ResolutionSuper-Resolution