paper-with-me

Papers

What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training

2025-06-01 · Marianne de Heer Kloots, Hosein Mohebbi, Charlotte Pouw, Gaofei Shen, Willem Zuidema, Martijn Bentum

How language-specific are speech representations learned by self-supervised models? Existing work has shown that a range of linguistic features can be successfully decoded from end-to-end models trained only on speech recordings. However, it's less clear to what extent pre-training on specific languages improves language-specific linguistic information. Here we test the encoding of Dutch phonetic and lexical information in internal representations of self-supervised Wav2Vec2 models. Pre-training exclusively on Dutch improves the representation of Dutch linguistic features as compared to pre-training on similar amounts of English or larger amounts of multilingual data. This language-specific advantage is well-detected by trained clustering or classification probes, and partially observable using zero-shot metrics. Furthermore, the language-specific benefit on linguistic feature encoding aligns with downstream performance on Automatic Speech Recognition.

📄 PDF Abstract BibTeX arXiv:2506.00981

Code (1)

mdhk/ssl-nl-eval 공식 구현

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

What Do Self-Supervised Speech Models Know About Words?

2023-06-30 · Ankita Pasad, Chung-Ming Chien, Shane Settle, Karen Livescu

Many self-supervised speech models (S3Ms) have been introduced over the last few years, improving performance and data efficiency on various speech tasks. However, these empirical successes alone do not give a complete p…

SentenceSentence SimilarityVisual Grounding

Probing self-attention in self-supervised speech models for cross-linguistic differences

2024-09-04 · Sai Gopinath, Joselyn Rodriguez

Speech models have gained traction thanks to increase in accuracy from novel transformer architectures. While this impressive increase in performance across automatic speech recognition (ASR) benchmarks is noteworthy, th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Speaker Group Encoding in Self-supervised Speech Recognition Models

2026-06-09 · Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre, François Portet arxiv

We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identification (SID), finetuned on automatic spee…

Speaker IdentificationSpeech Recognition

Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0

2024-07-03 · Marianne de Heer Kloots, Willem Zuidema

What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by…

BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation

2025-10-28 · Raphaël Bagat, Irina Illina, Emmanuel Vincent arxiv

Automatic Speech Recognition (ASR) systems, despite large multilingual training, struggle in low-resource scenarios where labeled data is scarce. We propose BEARD (BEST-RQ Encoder Adaptation with Re-training and Distilla…

Self-Supervised LearningKnowledge DistillationSpeech RecognitionDomain Adaptation