paper-with-me

홈 › Papers

GigaAM Multilingual: Foundation Model for Underrepresented Languages

2026-07-11 · Andrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin arxiv

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.

📄 PDF Abstract BibTeX arXiv:2607.10371

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM for Everyone: Representing the Underrepresented in Large Language Models

2024-09-20 · Samuel Cahyawijaya

Natural language processing (NLP) has witnessed a profound impact of large language models (LLMs) that excel in a multitude of tasks. However, the limitation of LLMs in multilingual settings, particularly in underreprese…

In-Context Learning

GigaAM: Efficient Self-Supervised Learner for Speech Recognition

2025-06-01 · Aleksandr Kutsakov, Alexandr Maximenko, Georgii Gospodinov, Pavel Bogomolov 외

Self-Supervised Learning (SSL) has demonstrated strong performance in speech processing, particularly in automatic speech recognition. In this paper, we explore an SSL pretraining framework that leverages masked language…

Automatic Speech RecognitionLanguage ModelingLanguage ModellingMasked Language Modeling+3

The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context

2025-04-03 · Nikhil Verma, Manasa Bharadwaj

Alignment tuning has enabled large language models to excel in reasoning, instruction-following, and minimizing harmful generations. However, despite their widespread deployment, these models exhibit a monolingual bias, …

Instruction Following

Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi

2026-03-03 · Shiza Fatimah, Aniket Sen, Sophia Falk, Florian Mai 외 arxiv

The dominance of large multilingual foundation models has widened linguistic inequalities in Natural Language Processing (NLP), often leaving low-resource languages underrepresented. This paper introduces LilMoo, a 0.6-b…

Continual Pretraining

Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages

2024-05-22 · Corinne Aars, Lauren Adams, Xiaokan Tian, Zhaoyu Wang 외

This study presents the development and evaluation of a ByT5-based multilingual translation model tailored for translating the Bible into underrepresented languages. Utilizing the comprehensive Johns Hopkins University B…