paper-with-me

Papers

Dual Script E2E framework for Multilingual and Code-Switching ASR

2021-06-02 · Mari Ganesh Kumar, Jom Kuriakose, Anand Thyagachandran, Arun Kumar A, Ashish Seth, Lodagala Durga Prasad, Saish Jaiswal, Anusha Prakash, Hema Murthy

India is home to multiple languages, and training automatic speech recognition (ASR) systems for languages is challenging. Over time, each language has adopted words from other languages, such as English, leading to code-mixing. Most Indian languages also have their own unique scripts, which poses a major limitation in training multilingual and code-switching ASR systems. Inspired by results in text-to-speech synthesis, in this work, we use an in-house rule-based phoneme-level common label set (CLS) representation to train multilingual and code-switching ASR for Indian languages. We propose two end-to-end (E2E) ASR systems. In the first system, the E2E model is trained on the CLS representation, and we use a novel data-driven back-end to recover the native language script. In the second system, we propose a modification to the E2E model, wherein the CLS representation and the native language characters are used simultaneously for training. We show our results on the multilingual and code-switching tasks of the Indic ASR Challenge 2021. Our best results achieve 6% and 5% improvement (approx) in word error rate over the baseline system for the multilingual and code-switching tasks, respectively, on the challenge development data.

📄 PDF Abstract BibTeX arXiv:2106.01400

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

Unsupervised Code-Switching for Multilingual Historical Document Transcription

2015-05-01 · HLT 2015 5 · Dan Klein, Taylor Berg-Kirkpatrick, Dan Garrette, Hannah Alpert-Abrams
Language IdentificationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1

Romanization Encoding For Multilingual ASR

2024-07-05 · Wen Ding, Fei Jia, Hainan Xu, Yu Xi 외

We introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated to…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

2026-05-13 · Trung Nguyen Quang, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He 외 arxiv

Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language…

Speech Recognition

Developing a Multilingual Dataset and Evaluation Metrics for Code-Switching: A Focus on Hong Kong's Polylingual Dynamics

2023-10-27 · Peng Xie, Kani Chen

The existing audio datasets are predominantly tailored towards single languages, overlooking the complex linguistic behaviors of multilingual communities that engage in code-switching. This practice, where individuals fr…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding

2024-06-17 · Haneul Yoo, Yongjin Yang, Hwaran Lee

As large language models (LLMs) have advanced rapidly, concerns regarding their safety have become prominent. In this paper, we discover that code-switching in red-teaming queries can effectively elicit undesirable behav…

16kLanguage ModellingRed TeamingSafety Alignment