paper-with-me

Papers

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

2025-11-20 · Wei-Cheng Tseng, David Harwath arxiv

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extractors for a broader range of speech processing tasks. Building on this trend, we introduce Codec2Vec, the first speech representation learning framework that relies exclusively on discrete audio codec units. This approach offers several advantages, including improved data storage and transmission efficiency, faster training, and enhanced data privacy. We explore masked prediction with various training target derivation strategies to thoroughly understand the effectiveness of this framework. Evaluated on the SUPERB benchmark, Codec2Vec achieves competitive performance compared to continuous-input models while reducing storage requirements by up to 16.5x and training time by 2.3x, showcasing its scalability and efficiency.

📄 PDF Abstract BibTeX arXiv:2511.16639

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeech Synthesis

Similar Papers 제목 키워드 기반

DM-Codec: Distilling Multimodal Representations for Speech Tokenization

2024-10-19 · Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman 외

Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete toke…

Self-Supervised LearningSpeech Tokenization

Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations

2025-03-15 · Xue Jiang, Xiulian Peng, Yuan Zhang, Yan Lu

Current large speech language models are mainly based on semantic tokens from discretization of self-supervised learned representations and acoustic tokens from a neural codec, following a semantic-modeling and acoustic-…

Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations

2024-07-03 · Kunal Dhawan, Nithin Rao Koluguri, Ante Jukić, Ryan Langman 외

Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker ver…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)QuantizationSpeaker Verification+2

FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

2025-09-14 · Md Mubtasim Ahasan, Rafat Hasan Khan, Tasnim Mohiuddin, Aman Chadha 외 arxiv

Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to …

Representation LearningSpeech Synthesis

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

2021-04-01 · Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov 외

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic inform…

DisentanglementRepresentation LearningResynthesisSpeaker Identification+1