paper-with-me

홈 › Papers

Language-specific Neurons Do Not Facilitate Cross-Lingual Transfer

2025-03-21 · Soumen Kumar Mondal, Sayambhu Sen, Abhishek Singhania, Preethi Jyothi

Multilingual large language models (LLMs) aim towards robust natural language understanding across diverse languages, yet their performance significantly degrades on low-resource languages. This work explores whether existing techniques to identify language-specific neurons can be leveraged to enhance cross-lingual task performance of lowresource languages. We conduct detailed experiments covering existing language-specific neuron identification techniques (such as Language Activation Probability Entropy and activation probability-based thresholding) and neuron-specific LoRA fine-tuning with models like Llama 3.1 and Mistral Nemo. We find that such neuron-specific interventions are insufficient to yield cross-lingual improvements on downstream tasks (XNLI, XQuAD) in lowresource languages. This study highlights the challenges in achieving cross-lingual generalization and provides critical insights for multilingual LLMs.

📄 PDF Abstract BibTeX arXiv:2503.17456

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferNatural Language Understanding

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

Linguistic Neuron Overlap Patterns to Facilitate Cross-lingual Transfer on Low-resource Languages

2025-08-23 · Yuemei Xu, Kexin Xu, Jian Zhou, Ling Hu 외 arxiv

The current Large Language Models (LLMs) face significant challenges in improving their performance on low-resource languages and urgently need data-efficient methods without costly fine-tuning. From the perspective of l…

Cross-Lingual Transfer

Isolating Culture Neurons in Multilingual Large Language Models

2025-08-04 · Danial Namazifard, Lukas Galke Poech arxiv

Language and culture are deeply intertwined, yet it has been unclear how and where multilingual large language models encode culture. Here, we build on an established methodology for identifying language-specific neurons…

The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs

2025-09-21 · Hinata Tezuka, Naoya Inoue arxiv

Recent studies have suggested a processing framework for multilingual inputs in decoder-based LLMs: early layers convert inputs into English-centric and language-agnostic representations; middle layers perform reasoning …

On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons

2024-04-03 · Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka 외

Current decoder-based pre-trained language models (PLMs) successfully demonstrate multilingual capabilities. However, it is unclear how these models handle multilingualism. We analyze the neuron-level internal behavior o…

DecoderText Generation

How do Large Language Models Handle Multilingualism?

2024-02-29 · Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi 외

Large language models (LLMs) have demonstrated impressive capabilities across diverse languages. This study explores how LLMs handle multilingualism. Based on observed language ratio shifts among layers and the relations…