paper-with-me

홈 › Papers

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

2024-08-29 · Shengpeng Ji, Ziyue Jiang, Wen Wang, Yifu Chen, Minghui Fang, Jialong Zuo, Qian Yang, Xize Cheng, Zehan Wang, RuiQi Li, Ziang Zhang, Xiaoda Yang, Rongjie Huang, Yidi Jiang, Qian Chen, Siqi Zheng, Zhou Zhao

Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1)extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2)improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The related code, demos, and pre-trained models are available at https://github.com/jishengpeng/WavTokenizer.

📄 PDF Abstract BibTeX arXiv:2408.16532

Code (1)

jishengpeng/wavtokenizer 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders

2025-10-27 · Nathan Paek, Yongyi Zang, Qihui Yang, Randal Leistikow arxiv

While sparse autoencoders (SAEs) successfully extract interpretable features from language models, applying them to audio generation faces unique challenges: audio's dense nature requires compression that obscures semant…

Audio GenerationMusic Generation

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

2026-06-01 · Hui Li, Yangfan Gao, Junlin Shang, Changhao Jiang 외 arxiv

Audio tokenizers serve as the discrete interface between continuous audio and Audio Language Models (ALMs), but existing tokenizers often struggle to support both understanding and generation. Reconstruction-oriented cod…

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer

2025-06-01 · Daiki Takeuchi, Binh Thien Nguyen, Masahiro Yasuda, Yasunori Ohishi 외

Automated Audio Captioning (AAC) aims to describe the semantic contexts of general sounds, including acoustic events and scenes, by leveraging effective acoustic features. To enhance performance, an AAC method, EnCLAP, e…

Audio captioningLanguage ModelingLanguage ModellingQuantization

Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer

2024-10-10 · Slava Shechtman, Avihu Dekel

Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. Various publicly available trainable disc…

Towards audio language modeling - an overview

2024-02-20 · Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 외

Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as suitable tokenizers for converting continu…

Language ModelingLanguage Modelling