paper-with-me

Papers

LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech

2023-09-11 · Titouan Parcollet, Ha Nguyen, Solene Evain, Marcely Zanon Boito, Adrien Pupier, Salima Mdhaffar, Hang Le, Sina Alisamir, Natalia Tomashenko, Marco Dinarelli, Shucong Zhang, Alexandre Allauzen, Maximin Coavoux, Yannick Esteve, Mickael Rouvier, Jerome Goulian, Benjamin Lecouteux, Francois Portet, Solange Rossato, Fabien Ringeval, Didier Schwab, Laurent Besacier

Self-supervised learning (SSL) is at the origin of unprecedented improvements in many different domains including computer vision and natural language processing. Speech processing drastically benefitted from SSL as most of the current domain-related tasks are now being approached with pre-trained models. This work introduces LeBenchmark 2.0 an open-source framework for assessing and building SSL-equipped French speech technologies. It includes documented, large-scale and heterogeneous corpora with up to 14,000 hours of heterogeneous speech, ten pre-trained SSL wav2vec 2.0 models containing from 26 million to one billion learnable parameters shared with the community, and an evaluation protocol made of six downstream tasks to complement existing benchmarks. LeBenchmark 2.0 also presents unique perspectives on pre-trained SSL models for speech with the investigation of frozen versus fine-tuned downstream models, task-agnostic versus task-specific pre-trained models as well as a discussion on the carbon footprint of large-scale model training. Overall, the newly introduced models trained on 14,000 hours of French speech outperform multilingual and previous LeBenchmark SSL models across the benchmark but also required up to four times more energy for pre-training.

📄 PDF Abstract BibTeX arXiv:2309.05472

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

LeBenchmark: A Reproducible Framework for Assessing Self-Supervised Representation Learning from Speech

2021-04-23 · Solene Evain, Ha Nguyen, Hang Le, Marcely Zanon Boito 외

Self-Supervised Learning (SSL) using huge unlabeled data has been successfully explored for image and natural language processing. Recent works also investigated SSL from speech. They were notably successful to improve p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionRepresentation Learning+5

Pantagruel: Unified Self-Supervised Encoders for French Text and Speech

2026-01-09 · Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain, Maryem Bouziane 외 arxiv

We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or speech units, Pantagruel learns contextua…

Representation Learning

What has LeBenchmark Learnt about French Syntax?

2024-03-04 · Zdravko Dugonjić, Adrien Pupier, Benjamin Lecouteux, Maximin Coavoux

The paper reports on a series of experiments aiming at probing LeBenchmark, a pretrained acoustic model trained on 7k hours of spoken French, for syntactic information. Pretrained acoustic models are increasingly used fo…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpoken Language Understanding

Deriving Benchmarking Datasets from Long-Form Recordings: Challenges and Opportunities

2026-07-03 · Kaveri K. Sheth, Lawrence Borst, Tarek Kunze, Marvin Lavechin 외 arxiv

Long-form recordings (LFRs) of child-centered audio are ecologically valid sources for studying early language development, but three problems limit their use. First, LFR corpora are collected across sites with heterogen…

Vers la compréhension automatique de la parole bout-en-bout à moindre effort

2022-07-01 · Marco Naguib, François Portet, Marco Dinarelli

Recent advances in spoken language understanding benefited from Self-Supervised models trained on large speech corpora. For French, the LeBenchmark project has made such models available and has led to impressive progres…

Spoken Language Understanding