FlowDec: A flow-based full-band general audio codec with high perceptual quality
We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from speech to general audio and move from 24 kbit/s to as low as 4 kbit/s, while improving output quality and reducing the required postfilter DNN evaluations from 60 to 6 without any fine-tuning or distillation techniques. We provide theoretical insights and geometric intuitions for our approach in comparison to ScoreDec as well as another recent work that uses flow matching, and conduct ablation studies on our proposed components. We show that FlowDec is a competitive alternative to the recent GAN-dominated stream of neural codecs, achieving FAD scores better than those of the established GAN-based codec DAC and listening test scores that are on par, and producing qualitatively more natural reconstructions for speech and harmonic structures in music.
Code (1)
Tasks
FADMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
FlowDec: Temporal Conditional Flow Decorruptor for Robust Continuous Vision-Language Navigation
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions in unseen scenes. While Large Models (LMs) have advanced VLN-CE, their performance remains severe…
Vision-Language NavigationImage RestorationFull-band General Audio Synthesis with Score-based Diffusion
Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limit…
Audio SynthesisDiversityRFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
Recent advancements in generative modeling have significantly enhanced the reconstruction of audio waveforms from various representations. While diffusion models are adept at this task, they are hindered by latency issue…
Audio GenerationComputational EfficiencyGPUSpeech SynthesisNative Multi-Band Audio Coding within Hyper-Autoencoded Reconstruction Propagation Networks
Spectral sub-bands do not portray the same perceptual relevance. In audio coding, it is therefore desirable to have independent control over each of the constituent bands so that bitrate assignment and signal reconstruct…
Bandwidth ExtensionPromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets an…
Response Generation