paper-with-me

Papers

Building Robust and Scalable Multilingual ASR for Indian Languages

2025-11-19 · Arjun Gangwar, Kaousheik Jayakumar, S. Umesh arxiv

This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR systems to improve in predicting the language and dialect of the utterance among 8 languages across 33 dialects. We participated in Track 1 and Track 2, which restricts the use of additional data and develop from-the-scratch multilingual systems. We presented a novel training approach using Multi-Decoder architecture with phonemic Common Label Set (CLS) as intermediate representation. It improved the performance over the baseline (in the CLS space). We also discuss various methods used to retain the gain obtained in the phonemic space while converting them back to the corresponding grapheme representations. Our systems beat the baseline in 3 languages (Track 2) in terms of WER/CER and achieved the highest language ID and dialect ID accuracy among all participating teams (Track 2).

📄 PDF Abstract BibTeX arXiv:2511.15418

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DuDe: Dual-Decoder Multilingual ASR for Indian Languages using Common Label Set

2022-10-30 · Arunkumar A, Mudit Batra, Umesh S

In a multilingual country like India, multilingual Automatic Speech Recognition (ASR) systems have much scope. Multilingual ASR systems exhibit many advantages like scalability, maintainability, and improved performance …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+2

SPRING-INX: A Multilingual Indian Language Speech Corpus by SPRING Lab, IIT Madras

2023-10-23 · Nithya R, Malavika S, Jordan F, Arjun Gangwar 외

India is home to a multitude of languages of which 22 languages are recognised by the Indian Constitution as official. Building speech based applications for the Indian population is a difficult problem owing to limited …

IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

2023-12-15 · Saiful Haq, Ashutosh Sharma, Pushpak Bhattacharyya

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major…

Information RetrievalMachine TranslationRetrieval

Multilingual and code-switching ASR challenges for low resource Indian languages

2021-04-01 · Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah, Ankita Singh 외

Recently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking advantage of low amounts of labeled corpora …

Automatic Speech Recognition (ASR)SentenceSpeech Recognition

Debiasing Multilingual Word Embeddings: A Case Study of Three Indian Languages

2021-07-21 · Srijan Bansal, Vishal Garimella, Ayush Suhane, Animesh Mukherjee

In this paper, we advance the current state-of-the-art method for debiasing monolingual word embeddings so as to generalize well in a multilingual setting. We consider different methods to quantify bias and different deb…

Multilingual Word EmbeddingsWord Embeddings