paper-with-me

Papers

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

2025-03-15 · Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-synthesis paradigm. However, semantic tokens discard paralinguistic attributes of speakers that is important for natural spoken communication, while prompt-based acoustic synthesis from semantic tokens has limits in recovering paralinguistic details and suffers from robustness issues, especially when there are domain gaps between the prompt and the target. This paper unifies two types of tokens and proposes the UniCodec, a universal speech token learning that encapsulates all semantics of speech, including linguistic and paralinguistic information, into a compact and semantically-disentangled unified token. Such a unified token can not only benefit speech language models in understanding with paralinguistic hints but also help speech generation with high-quality output. A low-bitrate neural codec is leveraged to learn such disentangled discrete representations at global and local scales, with knowledge distilled from self-supervised learned features. Extensive evaluations on multilingual datasets demonstrate its effectiveness in generating natural, expressive and long-term consistent output quality with paralinguistic attributes well preserved in several speech processing tasks.

📄 PDF Abstract BibTeX arXiv:2503.12115

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ultra-Low-Bitrate Speech Coding with Pretrained Transformers

2022-07-05 · Ali Siahkoohi, Michael Chinen, Tom Denton, W. Bastiaan Kleijn 외

Speech coding facilitates the transmission of speech over low-bandwidth networks with minimal distortion. Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional…

DecoderInductive Bias

LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

2024-10-21 · Yiwei Guo, Zhihan Li, Chenpeng Du, Hankun Wang 외

Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models. In this work, we p…

DisentanglementLanguage ModelingLanguage ModellingQuantization+1

FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks

2025-02-06 · Luca Della Libera, Francesco Paissan, Cem Subakan, Mirco Ravanelli

Large language models have revolutionized natural language processing through self-supervised pretraining on massive datasets. Inspired by this success, researchers have explored adapting these methods to speech by discr…

ResynthesisVoice Conversion

UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension

2025-05-22 · Kishan Gupta, Srikanth Korse, Andreas Brendel, Nicola Pia 외

In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, e…

Bandwidth ExtensionGenerative Adversarial Network

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

2024-04-03 · Yu Pan, Xiang Zhang, Yuguang Yang, Jixun Yao 외

Neural speech codecs have recently emerged as a focal point in the fields of speech compression and generation. Despite this progress, achieving high-quality speech reconstruction under low-bitrate scenarios remains a si…

Representation LearningSpeaker VerificationSpeech SynthesisSSIM+3