paper-with-me

홈 › Papers

Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models

2025-10-15 · Daniil Gurgurov, Tanja Baeumel, Josef van Genabith, Simon Ostermann arxiv

Large language models (LLMs) exhibit substantial performance disparities across languages, particularly between high- and low-resource settings. We propose a framework for improving performance in underrepresented languages while preserving general-purpose capabilities via targeted fine-tuning of sparse, language-associated subnetworks. Our approach identifies language-relevant neurons using Language Activation Probability Entropy (LAPE), an information-theoretic metric that reliably captures language-specific activation patterns, and fine-tunes only the corresponding weights. Experiments on Llama-3.1-8B, Mistral-Nemo-12B, and Aya-Expanse-8B across 12 mid- and low-resource languages show that our method consistently outperforms full fine-tuning, FFN-only fine-tuning, LoRA, IA^3, and random-subset baselines while updating only 0.2-1% of model parameters. We further show that sparse, neuron-targeted fine-tuning can inject new language capabilities without catastrophic forgetting, with potential applicability to other model capabilities. Mechanistic analyses of weight updates and internal representations reveal asymmetric roles of FFN projections in language adaptation and improved cross-lingual alignment. Finally, we release language neuron sets for over 100 languages together with our adaptation pipeline, enabling a cost-effective path for extending LLMs to underrepresented languages.

📄 PDF Abstract BibTeX arXiv:2510.13580

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation

2026-04-04 · Kening Zheng, Wei-Chieh Huang, Jiahao Huo, Zhonghao Li 외 arxiv

Mixture-of-Experts (MoE) models exhibit striking performance disparities across languages, yet the internal mechanisms driving these gaps remain poorly understood. In this work, we conduct a systematic analysis of expert…

Efficacy of ByT5 in Multilingual Translation of Biblical Texts for Underrepresented Languages

2024-05-22 · Corinne Aars, Lauren Adams, Xiaokan Tian, Zhaoyu Wang 외

This study presents the development and evaluation of a ByT5-based multilingual translation model tailored for translating the Bible into underrepresented languages. Utilizing the comprehensive Johns Hopkins University B…

LLM for Everyone: Representing the Underrepresented in Large Language Models

2024-09-20 · Samuel Cahyawijaya

Natural language processing (NLP) has witnessed a profound impact of large language models (LLMs) that excel in a multitude of tasks. However, the limitation of LLMs in multilingual settings, particularly in underreprese…

In-Context Learning

PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition

2021-06-10 · NeurIPS 2021 12 · Cheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang 외

Self-supervised speech representation learning (speech SSL) has demonstrated the benefit of scale in learning rich representations for Automatic Speech Recognition (ASR) with limited paired data, such as wav2vec 2.0. We …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+3

Evaluating Lottery Tickets Under Distributional Shifts

2019-10-28 · WS 2019 11 · Shrey Desai, Hongyuan Zhan, Ahmed Aly

The Lottery Ticket Hypothesis suggests large, over-parameterized neural networks consist of small, sparse subnetworks that can be trained in isolation to reach a similar (or better) test accuracy. However, the initializa…

Inductive Bias