paper-with-me

Papers

Variable Bitrate Residual Vector Quantization for Audio Coding

2024-10-08 · Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo, Yukara Ikemiya, Zhi Zhong, Kin Wai Cheuk, Marco A. Martínez-Ramírez, Kyogu Lee, Wei-Hsiang Liao, Yuki Mitsufuji

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple input audio, such as silence. To address this limitation, we propose variable bitrate RVQ (VRVQ) for audio codecs, which allows for more efficient coding by adapting the number of codebooks used per frame. Furthermore, we propose a gradient estimation method for the non-differentiable masking operation that transforms from the importance map to the binary importance mask, improving model training via a straight-through estimator. We demonstrate that the proposed training framework achieves superior results compared to the baseline method and shows further improvement when applied to the current state-of-the-art codec.

📄 PDF Abstract BibTeX arXiv:2410.06016

Code (0)

등록된 구현이 없습니다.

Tasks

Audio CompressionQuantization

Similar Papers 제목 키워드 기반

Switchcodec: Adaptive residual-expert sparse quantization for high-fidelity neural audio coding

2026-01-28 · Xiangbo Wang, Wenbin Jiang, Jin Wang, Yubo You 외 arxiv

Recent neural audio compression models often rely on residual vector quantization for high-fidelity coding, but using a fixed number of per-frame codebooks is suboptimal for the wide variability of audio content-especial…

Cross-Scale Vector Quantization for Scalable Neural Speech Coding

2022-07-07 · Xue Jiang, Xiulian Peng, Huaying Xue, Yuan Zhang 외

Bitrate scalability is a desirable feature for audio coding in real-time communications. Existing neural audio codecs usually enforce a specific bitrate during training, so different models need to be trained for each ta…

Quantization

SNAC: Multi-Scale Neural Audio Codec

2024-10-18 · Hubert Siuzdak, Florian Grötschla, Luca A. Lanzendörfer

Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and unders…

Audio CompressionAudio GenerationLanguage ModelingLanguage Modelling+1

ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs

2024-10-16 · Rui-Chen Zheng, Hui-Peng Du, Xiao-Hang Jiang, Yang Ai 외

Current neural audio codecs typically use residual vector quantization (RVQ) to discretize speech signals. However, they often experience codebook collapse, which reduces the effective codebook size and leads to suboptim…

DiversityOnline ClusteringQuantizationtext-to-speech+1

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

2025-07-17 · Anastasia Kuznetsova, Inseon Jang, Wootaek Lim, Minje Kim

Neural audio codecs, leveraging quantization algorithms, have significantly impacted various speech/audio tasks. While high-fidelity reconstruction is paramount for human perception, audio coding for machines (ACoM) prio…

Audio ClassificationAutomatic Speech RecognitionQuantizationspeech-recognition+1