paper-with-me

Papers

The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

2020-11-23 · Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé, Morgane Rivière, Evgeny Kharitonov, Alexei Baevski, Ewan Dunbar, Emmanuel Dupoux

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot metrics probing for the quality of the learned models at 4 linguistic levels: phonetics, lexicon, syntax and semantics. We present the results and analyses of a composite baseline made of the concatenation of three unsupervised systems: self-supervised contrastive representation learning (CPC), clustering (k-means) and language modeling (LSTM or BERT). The language models learn on the basis of the pseudo-text derived from clustering the learned representations. This simple pipeline shows better than chance performance on all four metrics, demonstrating the feasibility of spoken language modeling from raw speech. It also yields worse performance compared to text-based 'topline' systems trained on the same data, delineating the space to be explored by more sophisticated end-to-end models.

📄 PDF Abstract BibTeX arXiv:2011.11588

Code (2)

bootphon/zerospeech2021_baseline 공식 구현 pytorch
zerospeech/zerospeech2021_baseline pytorch

Tasks

ClusteringLanguage ModelingLanguage ModellingRepresentation Learning

Similar Papers 제목 키워드 기반

The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units

2020-10-12 · Ewan Dunbar, Julien Karadayi, Mathieu Bernard, Xuan-Nga Cao 외

We present the Zero Resource Speech Challenge 2020, which aims at learning speech representations from raw audio signals without any labels. It combines the data sets and metrics from two previous benchmarks (2017 and 20…

Speech Synthesis

The Zero Resource Speech Challenge 2017

2017-12-12 · Ewan Dunbar, Xuan Nga Cao, Juan Benjumea, Julien Karadayi 외

We describe a new challenge aimed at discovering subword and word units from raw speech. This challenge is the followup to the Zero Resource Speech Challenge 2015. It aims at constructing systems that generalize across l…

RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks

2026-03-02 · Alexandra Diaconu, Mădălina Vînaga, Bogdan Alexe arxiv

We introduce RO-N3WS, a benchmark Romanian speech dataset designed to improve generalization in automatic speech recognition (ASR), particularly in low-resource and out-of-distribution (OOD) conditions. RO-N3WS comprises…

Speech RecognitionDomain Adaptation

Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge

2022-10-27 · Ewan Dunbar, Nicolas Hamilakis, Emmanuel Dupoux

Recent progress in self-supervised or unsupervised machine learning has opened the possibility of building a full speech processing system from raw audio without using any textual representations or expert labels such as…

Acoustic Unit DiscoveryLanguage ModelingLanguage ModellingResynthesis

Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages

2023-10-04 · Kuan-Po Huang, Chih-Kai Yang, Yu-Kuan Fu, Ewan Dunbar 외

We introduce a new zero resource code-switched speech benchmark designed to directly assess the code-switching capabilities of self-supervised speech encoders. We showcase a baseline system of language modeling on discre…

Language ModelingLanguage Modelling