paper-with-me

Papers

Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

2024-09-21 · Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, KaiWei Chang, Jiawei Du, Ke-Han Lu, Alexander H. Liu, Ho-Lam Chung, Yuan-Kuei Wu, Dongchao Yang, Songxiang Liu, Yi-Chiao Wu, Xu Tan, James Glass, Shinji Watanabe, Hung-Yi Lee

Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 2024, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge's rules, datasets, five participant systems, results, and findings.

📄 PDF Abstract BibTeX arXiv:2409.14085

Code (1)

ga642381/speech-trident 공식 구현 tf

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

2024-02-20 · Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin, Yuan-Kuei Wu 외

The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance. Recent years have witnessed significant developments in codec models. The ideal sound cod…

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

2025-11-20 · Wei-Cheng Tseng, David Harwath arxiv

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extra…

Representation LearningSpeech Synthesis

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

2025-09-15 · Adhiraj Banerjee, Vipul Arora arxiv

Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-latency edge or codec-mediated deployment. Ex…

One Quantizer is Enough: Toward a Lightweight Audio Codec

2025-04-07 · Linwei Zhai, Han Ding, Cui Zhao, Fei Wang 외

Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and generate discrete tokens that can be utilized in downstream generative modeling tasks. However, leading approaches o…

HILCodec: High-Fidelity and Lightweight Neural Audio Codec

2024-05-08 · Sunghwan Ahn, Beom Jun Woo, Min Hyun Han, Chanyeong Moon 외

The recent advancement of end-to-end neural audio codecs enables compressing audio at very low bitrates while reconstructing the output audio with high fidelity. Nonetheless, such improvements often come at the cost of i…