paper-with-me

홈 › Papers

CLSRIL-23: Cross Lingual Speech Representations for Indic Languages

2021-07-15 · Anirudh Gupta, Harveen Singh Chadha, Priyanshi Shah, Neeraj Chhimwal, Ankur Dhuriya, Rishabh Gaur, Vivek Raghavan

We present a CLSRIL-23, a self supervised learning based audio pre-trained model which learns cross lingual speech representations from raw audio across 23 Indic languages. It is built on top of wav2vec 2.0 which is solved by training a contrastive task over masked latent speech representations and jointly learns the quantization of latents shared across all languages. We compare the language wise loss during pretraining to compare effects of monolingual and multilingual pretraining. Performance on some downstream fine-tuning tasks for speech recognition is also compared and our experiments show that multilingual pretraining outperforms monolingual training, in terms of learning speech representations which encodes phonetic similarity of languages and also in terms of performance on down stream tasks. A decrease of 5% is observed in WER and 9.5% in CER when a multilingual pretrained model is used for finetuning in Hindi. All the code models are also open sourced. CLSRIL-23 is a model trained on $23$ languages and almost 10,000 hours of audio data to facilitate research in speech recognition for Indic languages. We hope that new state of the art systems will be created using the self supervised approach, especially for low resources Indic languages.

📄 PDF Abstract BibTeX arXiv:2107.07402

Code (2)

Open-Speech-EkStep/vakyansh-models 공식 구현
Open-Speech-EkStep/vakyansh-wav2vec2-experimentation 공식 구현 pytorch

Tasks

Self-Supervised Learningspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Few-Shot Contrastive Adaptation for Audio Abuse Detection in Low-Resource Indic Languages

2026-04-10 · Aditya Narayan Sankaran, Reza Farahbakhsh, Noel Crespi arxiv

Abusive speech detection is becoming increasingly important as social media shifts towards voice-based interaction, particularly in multilingual and low-resource settings. Most current systems rely on automatic speech re…

Speech Recognition

Limitations of Cross-Lingual Learning from Image Search

2017-09-18 · WS 2018 7 · Mareike Hartmann, Anders Soegaard

Cross-lingual representation learning is an important step in making NLP scale to all the world's languages. Recent work on bilingual lexicon induction suggests that it is possible to learn cross-lingual representations …

Bilingual Lexicon InductionImage RetrievalRepresentation LearningTranslation

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison

2025-01-04 · Tsz Kin Lam, Marco Gaido, Sara Papi, Luisa Bentivogli 외

Following the remarkable success of Large Language Models (LLMs) in NLP tasks, there is increasing interest in extending their capabilities to speech -- the most common form of communication. The most widespread approach…

DecoderKnowledge DistillationSpeech-to-Text

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations

2026-06-24 · Ram Annamdevula, Ankit Tatawat, Ashishkumar P. Gudmalwar, Nirmesh J. Shah 외 arxiv

Accent conversion and controllability remain fundamental challenges in cross-lingual text-to-speech (TTS), particularly for low-resource and phonetically diverse Indic languages. While recent large language model (LLM)-b…

Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease

2026-03-23 · Abner Hernandez, Eunjung Yeo, Kwanghee Choi, Chin-Jou Li 외 arxiv

The limited availability of dysarthric speech data makes cross-lingual detection an important but challenging problem. A key difficulty is that speech representations often encode language-dependent structure that can co…