paper-with-me

홈 › Papers

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

2026-04-03 · Baban Gain, Asif Ekbal, Trilok Nath Singh arxiv

Weight-space model merging combines independently fine-tuned models without accessing original training data, offering a practical alternative to joint training. While merging succeeds in multitask settings, its behavior in multilingual contexts remains poorly understood. We systematically study weight-space merging for multilingual machine translation by fully fine-tuning language model on large-scale bilingual corpora and evaluating standard merging strategies. Our experiments reveal that merging degrades performance, especially when target languages differ. To explain this failure, we analyze internal representations using span-conditioned neuron selectivity and layer-wise centered kernel alignment. We find that language-specific neurons concentrate in embedding layers and upper transformer blocks, while intermediate layers remain largely shared across languages. Critically, fine-tuning redistributes rather than sharpens language selectivity: neurons for supervised and related languages become less exclusive, while those for unsupervised languages grow more isolated. This redistribution increases representational divergence in higher layers that govern generation. These findings suggest that multilingual fine-tuning may reshape geometry in ways that reduce compatibility with standard weight-space merging assumptions. Our work thus provides an explanation for why merging fails in multilingual translation scenarios.

📄 PDF Abstract BibTeX arXiv:2604.02881

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Do Multilingual Language Models Think Better in English?

2023-08-02 · Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle 외

Translate-test is a popular technique to improve the performance of multilingual language models. This approach works by translating the input into English using an external machine translation system, and running infere…

Common Sense ReasoningCross-Lingual Natural Language InferenceCross-Lingual Paraphrase IdentificationLanguage Modeling+6

Deep Reinforcement Learning From Raw Pixels in Doom

2016-10-07 · Danijar Hafner

Using current reinforcement learning methods, it has recently become possible to learn to play unknown 3D games from raw pixels. In this work, we study the challenges that arise in such complex environments, and summariz…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Resource Creation and Evaluation for Multilingual Sentiment Analysis in Social Media Texts

2014-05-01 · LREC 2014 5 · Alex Balahur, ra, Marco Turchi, Ralf Steinberger 외

This paper presents an evaluation of the use of machine translation to obtain and employ data for training multilingual sentiment classifiers. We show that the use of machine translated data obtained similar results as t…

ClassificationGeneral ClassificationMachine TranslationNatural Language Inference+4

A Multilingual Neural Machine Translation Model for Biomedical Data

2020-08-06 · EMNLP (NLP-COVID19) 2020 12 · Alexandre Bérard, Zae Myung Kim, Vassilina Nikoulina, Eunjeong L. Park 외

We release a multilingual neural machine translation model, which can be used to translate text in the biomedical domain. The model can translate from 5 languages (French, German, Italian, Korean and Spanish) into Englis…

Machine TranslationTranslation

Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language

2024-10-31 · Jiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin 외

English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs still underperform for non-English langua…