Does Transliteration Help Multilingual Language Modeling?
Script diversity presents a challenge to Multilingual Language Models (MLLM) by reducing lexical overlap among closely related languages. Therefore, transliterating closely related languages that use different writing scripts to a common script may improve the downstream task performance of MLLMs. We empirically measure the effect of transliteration on MLLMs in this context. We specifically focus on the Indic languages, which have the highest script diversity in the world, and we evaluate our models on the IndicGLUE benchmark. We perform the Mann-Whitney U test to rigorously verify whether the effect of transliteration is significant or not. We find that transliteration benefits the low-resource languages without negatively affecting the comparatively high-resource languages. We also measure the cross-lingual representation similarity of the models using centered kernel alignment on parallel sentences from the FLORES-101 dataset. We find that for parallel sentences across different languages, the transliteration-based model learns sentence representations that are more similar.
Code (1)
Tasks
DiversityLanguage ModelingLanguage ModellingMultiple Choice Question Answering (MCQA)Named Entity Recognition (NER)News ClassificationSentenceSentiment AnalysisTransliterationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
How Transliterations Improve Crosslingual Alignment
Recent studies have shown that post-aligning multilingual pretrained language models (mPLMs) using alignment objectives on both original and transliterated data can improve crosslingual alignment. This improvement furthe…
SentenceTransliterationInvestigating Lexical Sharing in Multilingual Machine Translation for Indian Languages
Multilingual language models have shown impressive cross-lingual transfer ability across a diverse set of languages and tasks. To improve the cross-lingual ability of these models, some strategies include transliteration…
Cross-Lingual TransferMachine TranslationTranslationTransliterationA Large-scale Evaluation of Neural Machine Transliteration for Indic Languages
We take up the task of large-scale evaluation of neural machine transliteration between English and Indic languages, with a focus on multilingual transliteration to utilize orthographic similarity between Indian language…
TranslationTransliterationLeveraging Orthographic Similarity for Multilingual Neural Transliteration
We address the task of joint training of transliteration models for multiple language pairs (multilingual transliteration). This is an instance of multitask learning, where individual tasks (language pairs) benefit from …
DecoderInformation RetrievalMachine TranslationMulti-Task Learning+1Language-agnostic Multilingual Modeling
Multilingual Automated Speech Recognition (ASR) systems allow for the joint training of data-rich and data-scarce languages in a single model. This enables data and parameter sharing across languages, which is especially…
speech-recognitionSpeech RecognitionTransliteration