paper-with-me

Papers

Using Random Codebooks for Audio Neural AutoEncoders

2024-09-25 · Benoît Giniès, Xiaoyu Bie, Olivier Fercoq, Gaël Richard

Latent representation learning has been an active field of study for decades in numerous applications. Inspired among others by the tokenization from Natural Language Processing and motivated by the research of a simple data representation, recent works have introduced a quantization step into the feature extraction. In this work, we propose a novel strategy to build the neural discrete representation by means of random codebooks. These codebooks are obtained by randomly sampling a large, predefined fixed codebook. We experimentally show the merits and potential of our approach in a task of audio compression and reconstruction.

📄 PDF Abstract BibTeX arXiv:2409.16677

Code (0)

등록된 구현이 없습니다.

Tasks

Audio CompressionQuantizationRepresentation Learning

Similar Papers 제목 키워드 기반

Learning Product Codebooks using Vector Quantized Autoencoders for Image Retrieval

2018-07-12 · Hanwei Wu, Markus Flierl

Vector-Quantized Variational Autoencoders (VQ-VAE)[1] provide an unsupervised model for learning discrete representations by combining vector quantization and autoencoders. In this paper, we study the use of VQ-VAE for r…

Image RetrievalQuantizationRepresentation LearningRetrieval

Semantic Codebooks as Effective Priors for Neural Speech Compression

2025-12-25 · Liuyang Bai, Weiyi Lu, Li Guo arxiv

Speech codecs are traditionally optimized for waveform fidelity, allocating bits to preserve acoustic detail even when much of it can be inferred from linguistic structure. This leads to inefficient compression and subop…

Variable Bitrate Residual Vector Quantization for Audio Coding

2024-10-08 · Yunkee Chae, Woosung Choi, Yuhta Takida, Junghyun Koo 외

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be subopti…

Audio CompressionQuantization

Probing neural audio codecs for distinctions among English nuclear tunes

2026-03-14 · Juan Pablo Vigneaux, Jennifer Cole arxiv

State-of-the-art spoken dialogue models (Défossez et al. 2024; Schalkwyk et al. 2025) use neural audio codecs to "tokenize" audio signals into a lower-frequency stream of vectorial latent representations, each quantized …

Enhancing Suno's Bark Text-to-Speech Model: Addressing Limitations Through Meta's Encodec and Pre-Trained Hubert

2023-04-18 · Social Science Research Network (SSRN) 2023 4 · Devin Schumacher, Francis LaBounty Jr.

Bark, a transformer-based text-to-audio model by Suno, generates highly realistic, multilingual speech as well as other audio, including music, background noise, and simple sound effects. While this model has shown promi…

Audio GenerationExpressive Speech SynthesisSpeech Synthesistext-to-speech+3