Papers Audio Compression
“Audio Compression” 태그가 달린 논문 42편 · 필터 해제
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. …
Audio CompressionLearning to Upsample and Upmix Audio in the Latent Domain
Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-to…
Audio CompressionBandwidth ExtensionComputational EfficiencyLAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
This paper introduces LAV (Latent Audio-Visual), a system that integrates EnCodec's neural audio compression with StyleGAN2's generative capabilities to produce visually dynamic outputs driven by pre-recorded audio. Unli…
Audio CompressionAudio Compression using Periodic Gabor with Biorthogonal Exchange: Implementation Using the Zak Transform
An efficient new approach to signal compression is presented based of a novel variation on the Gabor basis set. Following earlier work by Shimshovitz and Tannor, we convolve the conventional Gabor functions with Dirichle…
Audio CompressionCPUMusic2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoenc…
Audio CompressionDenoisingInformation RetrievalMusic Information Retrieval+1SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
In this paper, we propose "SoundSpring", a cutting-edge error-resilient audio transceiver that marries the robustness benefits of joint source-channel coding (JSCC) while also being compatible with current digital commun…
Audio CompressionLanguage ModelingLanguage ModellingMasked Language Modeling+1Optimizing Audio Compression Through Entropy-Controlled Dithering
This paper explores entropy-controlled dithering techniques in audio compression, examining the application of standard and modified TPDFs, combined with noise shaping and entropy-controlled parameters, across various au…
Audio CompressionRhythmTS3-Codec: Transformer-Based Simple Streaming Single Codec
Neural audio codecs (NACs) have garnered significant attention as key technologies for audio compression as well as audio representation for speech language models. While mainstream NAC models are predominantly convoluti…
Audio CompressionResidual vector quantization for KV cache compression in large language model
KV cache compression methods have mainly relied on scalar quantization techniques to reduce the memory requirements during decoding. In this work, we apply residual vector quantization, which has been widely used for hig…
Audio CompressionLanguage ModelingLanguage ModellingLarge Language Model+1SNAC: Multi-Scale Neural Audio Codec
Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and unders…
Audio CompressionAudio GenerationLanguage ModelingLanguage Modelling+1Variable Bitrate Residual Vector Quantization for Audio Coding
Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be subopti…
Audio CompressionQuantizationFlowMAC: Conditional Flow Matching for Audio Coding at Low Bit Rates
This paper introduces FlowMAC, a novel neural audio codec for high-quality general audio compression at low bit rates based on conditional flow matching (CFM). FlowMAC jointly learns a mel spectrogram encoder, quantizer …
Audio CompressionCPUDecoderUsing Random Codebooks for Audio Neural AutoEncoders
Latent representation learning has been an active field of study for decades in numerous applications. Inspired among others by the tokenization from Natural Language Processing and motivated by the research of a simple …
Audio CompressionQuantizationRepresentation LearningNDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
Built upon vector quantization (VQ), discrete audio codec models have achieved great success in audio compression and auto-regressive audio generation. However, existing models face substantial challenges in perceptual q…
Audio CompressionAudio GenerationDecoderQuantization+1Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, aud…
Audio CompressionLanguage ModelingLanguage ModellingQuantization+2Learning Source Disentanglement in Neural Audio Codec
Neural audio codecs have significantly advanced audio compression by efficiently converting continuous audio signals into discrete tokens. These codecs preserve high-quality sound and enable sophisticated sound generatio…
Audio CompressionAudio GenerationDisentanglementResynthesisCodec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
Recent advancements in audio generation have been significantly propelled by the capabilities of Large Language Models (LLMs). The existing research on audio LLM has primarily focused on enhancing the architecture and sc…
Audio CompressionAudio GenerationLanguage ModelingLanguage Modelling+4Music2Latent: Consistency Autoencoders for Latent Audio Compression
Efficient audio representations in a compressed continuous latent space are critical for generative audio modeling and Music Information Retrieval (MIR) tasks. However, some existing audio autoencoders have limitations, …
Audio CompressionInformation RetrievalMusic Information RetrievalAutoregressive Speech Synthesis without Vector Quantization
We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypa…
Audio CompressionDiversityLanguage ModelingLanguage Modelling+6Gull: A Generative Multifunctional Audio Codec
We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time com…
Audio CompressionAudio Source SeparationAudio Super-ResolutionDecoder+2