paper-with-me

홈 › Papers

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

2024-09-16 · Li-Wei Chen, Takuya Higuchi, He Bai, Ahmed Hussen Abdelaziz, Alexander Rudnicky, Shinji Watanabe, Tatiana Likhomanenko, Barry-John Theobald, Zakaria Aldeneh

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmasked context. The choice of prediction targets in this framework impacts their performance on downstream tasks. For instance, models pre-trained with targets that capture prosody learn representations suited for speaker-related tasks, while those pre-trained with targets that capture phonetics learn representations suited for content-related tasks. Moreover, prediction targets can differ in the level of detail they capture. Models pre-trained with targets that encode fine-grained acoustic features perform better on tasks like denoising, while those pre-trained with targets focused on higher-level abstractions are more effective for content-related tasks. Despite the importance of prediction targets, the design choices that affect them have not been thoroughly studied. This work explores the design choices and their impact on downstream task performance. Our results indicate that the commonly used design choices for HuBERT can be suboptimal. We propose approaches to create more informative prediction targets and demonstrate their effectiveness through improvements across various downstream tasks.

📄 PDF Abstract BibTeX arXiv:2409.10788

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingPrediction

Similar Papers 제목 키워드 기반

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

2026-06-29 · Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini arxiv

Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a…

Self-Supervised Learning

Alethia: A Foundational Encoder for Voice Deepfakes

2026-04-30 · Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya Koppisetti arxiv

Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In …

Zero-shot GeneralizationDeepFake Detection

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

2025-11-20 · Wei-Cheng Tseng, David Harwath arxiv

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extra…

Representation LearningSpeech Synthesis

Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation

2023-06-15 · Ziyang Ma, Zhisheng Zheng, Guanrou Yang, Yu Wang 외

The excellent generalization ability of self-supervised learning (SSL) for speech foundation models has garnered significant attention. HuBERT is a successful example that utilizes offline clustering to convert speech fe…

Automatic Speech RecognitionClusteringLanguage ModelingLanguage Modelling+3

Jointly Learning Visual and Auditory Speech Representations from Raw Data

2022-12-12 · Alexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 외

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1