paper-with-me

Papers

Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads

2023-08-28 · Salah Zaiem, Youcef Kemiche, Titouan Parcollet, Slim Essid, Mirco Ravanelli

Self-supervised learning (SSL) leverages large datasets of unlabeled speech to reach impressive performance with reduced amounts of annotated data. The high number of proposed approaches fostered the emergence of comprehensive benchmarks that evaluate their performance on a set of downstream tasks exploring various aspects of the speech signal. However, while the number of considered tasks has been growing, most proposals rely upon a single downstream architecture that maps the frozen SSL representations to the task labels. This study examines how benchmarking results are affected by changes in the probing head architecture. Interestingly, we found that altering the downstream architecture structure leads to significant fluctuations in the performance ranking of the evaluated models. Against common practices in speech SSL benchmarking, we evaluate larger-capacity probing heads, showing their impact on performance, inference costs, generalization and multi-level feature exploitation.

📄 PDF Abstract BibTeX arXiv:2308.14456

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?

2023-06-01 · Salah Zaiem, Youcef Kemiche, Titouan Parcollet, Slim Essid 외

Self-supervised learning (SSL) has recently allowed leveraging large datasets of unlabeled speech signals to reach impressive performance on speech tasks using only small amounts of annotated data. The high number of pro…

BenchmarkingDecoderSelf-Supervised Learning

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

2026-06-24 · Shu Shang, Fuliang Weng, Zeqian Hu, Yaqian Zhou arxiv

While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic representations behave under fine-grained dialect variation. Existing…

Self-Supervised Speech Representation Learning: A Review

2022-05-21 · Abdelrahman Mohamed, Hung-Yi Lee, Lasse Borgholt, Jakob D. Havtorn 외

Although supervised deep learning has revolutionized speech and audio processing, it has necessitated the building of specialist models for individual tasks and application scenarios. It is likewise difficult to apply th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingRepresentation Learning+3

S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations

2021-10-12 · Wen-Chin Huang, Shu-wen Yang, Tomoki Hayashi, Hung-Yi Lee 외

This paper introduces S3PRL-VC, an open-source voice conversion (VC) framework based on the S3PRL toolkit. In the context of recognition-synthesis VC, self-supervised speech representation (S3R) is valuable in its potent…

BenchmarkingVoice Conversion

BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models

2023-06-02 · Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón 외

Self-supervised techniques for learning speech representations have been shown to develop linguistic competence from exposure to speech without the need for human labels. In order to fully realize the potential of these …

BenchmarkingLanguage Acquisition