paper-with-me

홈 › Papers

What Can Low Resource Languages Learn From Each Other?

2026-08-27 · Achyuth P, Kahaan Shah, Chetan Arora arxiv

Despite the rapid advancement of Vision-Language Models (VLMs), their linguistic reach remains largely confined to high-resource languages, leaving the majority of the world's 7,000+ living languages on the wrong side of a growing digital divide. This disparity is especially pronounced in Optical Character Recognition (OCR), where low-resource scripts lack the massive datasets required for traditional scaling laws. We investigate OCR adaptation in extreme data-scarce regimes (<10K real and <250K synthetic images), demonstrating that conventional fine-tuning strategies often reach a performance ceiling. Our key finding reveals a structural inefficiency in language-specific adaptation: while higher layers of specialized models diverge to capture unique script nuances, the lower layers learn redundant, highly similar features. Motivated by this observation, we propose PSMC (Pre-train, Specialize, Merge, and Co-train), a data-efficient framework that capitalizes on a cross-script "transfer effect". Our approach first derives language-specific experts from a high-resource base model, then employs task arithmetic to fuse these experts into a unified, high-performance multilingual back- bone. Extensive evaluation across 10 Indian scripts (supporting 20+ languages) shows that PSMC achieves a ~2% average improvement in Word Recognition Rate (WRR) over individual specialist models without increasing parameter count. Our results indicate that joint training in the merged latent space facilitates a constructive knowledge transfer that benefits all constituent scripts, providing a scalable pathway for inclusive VLM development. Source code and datasets will be released post publication.

📄 PDF Abstract BibTeX arXiv:2608.27753

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What a Creole Wants, What a Creole Needs

2022-06-01 · LREC 2022 6 · Heather Lent, Kelechi Ogueji, Miryam de Lhoneux, Orevaoghene Ahia 외

In recent years, the natural language processing (NLP) community has given increased attention to the disparity of efforts directed towards high-resource languages over low-resource ones. Efforts to remedy this delta oft…

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages

2025-05-27 · Praveen Srinivasa Varadhan, Srija Anand, Soma Siddhartha, Mitesh M. Khapra

What happens when an English Fairytaler is fine-tuned on Indian languages? We evaluate how the English F5-TTS model adapts to 11 Indian languages, measuring polyglot fluency, voice-cloning, style-cloning, and code-mixing…

Synthetic Data GenerationVoice Cloning

Zero-Resource Multilingual Model Transfer: Learning What to Share

2018-09-27 · Xilun Chen, Ahmed Hassan Awadallah, Hany Hassan, Wei Wang 외

Modern natural language processing and understanding applications have enjoyed a great boost utilizing neural networks models. However, this is not the case for most languages especially low-resource ones with insufficie…

Cross-Lingual TransferMixture-of-Expertstext-classificationText Classification+1

WhatsApp Tiplines and Multilingual Claims in the 2021 Indian Assembly Elections

2025-07-22 · Gautam Kishore Shahi, Scot A. Hale arxiv

WhatsApp tiplines, first launched in 2019 to combat misinformation, enable users to interact with fact-checkers to verify misleading content. This study analyzes 580 unique claims (tips) from 451 users, covering both hig…

Multi-Source Cross-Lingual Model Transfer: Learning What to Share

2018-10-08 · ACL 2019 7 · Xilun Chen, Ahmed Hassan Awadallah, Hany Hassan, Wei Wang 외

Modern NLP applications have enjoyed a great boost utilizing neural networks models. Such deep neural models, however, are not applicable to most human languages due to the lack of annotated training data for various NLP…

Cross-Lingual NERCross-Lingual TransferMixture-of-Expertstext-classification+2