paper-with-me

Papers

Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training

2022-06-21 · Chengyi Wang, Yiming Wang, Yu Wu, Sanyuan Chen, Jinyu Li, Shujie Liu, Furu Wei

Recently, masked prediction pre-training has seen remarkable progress in self-supervised learning (SSL) for speech recognition. It usually requires a codebook obtained in an unsupervised way, making it less accurate and difficult to interpret. We propose two supervision-guided codebook generation approaches to improve automatic speech recognition (ASR) performance and also the pre-training efficiency, either through decoding with a hybrid ASR system to generate phoneme-level alignments (named PBERT), or performing clustering on the supervised speech features extracted from an end-to-end CTC model (named CTC clustering). Both the hybrid and CTC models are trained on the same small amount of labeled speech as used in fine-tuning. Experiments demonstrate significant superiority of our methods to various SSL and self-training baselines, with up to 17.0% relative WER reduction. Our pre-trained models also show good transferability in a non-ASR speech task.

📄 PDF Abstract BibTeX arXiv:2206.10125

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringSelf-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation

2025-09-23 · Roy Fejgin, Paarth Neekhara, Xuesong Yang, Edresson Casanova 외 arxiv

Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook structure. At each timestep, models must…

Computational Efficiency

SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision

2025-12-23 · Maxime Poli, Mahi Luthra, Youssef Benchekroun, Yosuke Higuchi 외 arxiv

The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This requires extracting semantic representa…

Representation LearningOnline Clustering

Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

2024-02-19 · Shengpeng Ji, Minghui Fang, Jialong Zuo, Ziyue Jiang 외

In recent years, large language models have achieved significant success in generative tasks related to speech, audio, music, and other signal domains. A crucial element of these models is the discrete acoustic codecs, w…

Audio CompressionAudio GenerationQuantization

Semantic Codebooks as Effective Priors for Neural Speech Compression

2025-12-25 · Liuyang Bai, Weiyi Lu, Li Guo arxiv

Speech codecs are traditionally optimized for waveform fidelity, allocating bits to preserve acoustic detail even when much of it can be inferred from linguistic structure. This leads to inefficient compression and subop…

Accented Speech Recognition With Accent-specific Codebooks

2023-10-24 · Darshan Prabhu, Preethi Jyothi, Sriram Ganapathy, Vinit Unni

Speech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems. Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of AS…

Accented Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1