paper-with-me

Papers

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

2023-11-25 · Tolúlopé Ògúnrèmí, Christopher D. Manning, Dan Jurafsky

While many speakers of low-resource languages regularly code-switch between their languages and other regional languages or English, datasets of codeswitched speech are too small to train bespoke acoustic models from scratch or do language model rescoring. Here we propose finetuning self-supervised speech representations such as wav2vec 2.0 XLSR to recognize code-switched data. We find that finetuning self-supervised multilingual representations and augmenting them with n-gram language models trained from transcripts reduces absolute word error rates by up to 20% compared to baselines of hybrid models trained from scratch on code-switched data. Our findings suggest that in circumstances with limited training data finetuning self-supervised representations is a better performing and viable solution.

📄 PDF Abstract BibTeX arXiv:2311.15077

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

XLSR 설명 없음

Similar Papers 제목 키워드 기반

Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models

2021-10-07 · Liang-Hsuan Tseng, Yu-Kuan Fu, Heng-Jui Chang, Hung-Yi Lee

Code-switching (CS) is common in daily conversations where more than one language is used within a sentence. The difficulties of CS speech recognition lie in alternating languages and the lack of transcribed data. Theref…

Language IdentificationSelf-Supervised LearningSentencespeech-recognition+1

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

2025-12-22 · Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber 외 arxiv

This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on …

Self-Supervised LearningRepresentation Learning

Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease

2026-03-23 · Abner Hernandez, Eunjung Yeo, Kwanghee Choi, Chin-Jou Li 외 arxiv

The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can co…

MAESTRO: Matched Speech Text Representations through Modality Matching

2022-04-07 · Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran 외

We present Maestro, a self-supervised training method to unify representations learnt from speech and text modalities. Self-supervised learning from speech signals aims to learn the latent structure inherent in the signa…

Language ModellingSelf-Supervised Learningspeech-recognitionSpeech Recognition+1

Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information

2022-12-07 · Fenglin Ding, Genshun Wan, Pengcheng Li, Jia Pan 외

Multilingual end-to-end models have shown great improvement over monolingual systems. With the development of pre-training methods on speech, self-supervised multilingual speech representation learning like XLSR has show…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+2