Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations
A concept can carry different associations across languages, while modern language models learn English alongside many other languages during pretraining. Yet comparisons among existing models cannot easily isolate how any one language changes the way these models represent English concepts because their training corpora, compute, architectures, and random seeds all differ. We study this question through a controlled experiment with 40 matched 310M-parameter decoder-only models that share an architecture, tokenizer, training recipe, and English data source. Each bilingual condition adds one of eight languages, while four experimental comparisons separately account for English exposure, total training, and English-document overlap. We align each model pair using 3,000 common English words, then measure where 1,000 held-out English concepts fall along 50 fixed semantic contrasts, such as red versus white. Across 32 experimental comparisons, English concept positions differ more between bilingual and English-only conditions than between English-only runs with different random seeds. These differences are larger in contextual states than in token embeddings and peak in middle layers. The language learned alongside English can therefore change how a model represents English concepts even when its English input representations are explicitly aligned.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
Recent work demonstrates the potential of multilingual pretraining of creating one model that can be used for various tasks in different languages. Previous work in multilingual pretraining has demonstrated that machine …
Machine TranslationTranslationThe Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
Multilingual large language models achieve impressive cross-lingual performance despite largely monolingual pretraining. While bilingual data in pretraining corpora is widely believed to enable these abilities, details o…
Improving the Lexical Ability of Pretrained Language Models for Unsupervised Neural Machine Translation
Successful methods for unsupervised neural machine translation (UNMT) employ crosslingual pretraining via self-supervision, often in the form of a masked language modeling or a sequence generation task, which requires th…
Bilingual Lexicon InductionLanguage ModelingLanguage ModellingMachine Translation+2RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning
In recent studies, it has been shown that Multilingual language models underperform their monolingual counterparts. It is also a well-known fact that training and maintaining monolingual models for each language is a cos…
Cross-Lingual TransferTransfer LearningSurprisal Predicts Code-Switching in Chinese-English Bilingual Text
Why do bilinguals switch languages within a sentence? The present observational study asks whether word surprisal and word entropy predict code-switching in bilingual written conversation. We describe and model a new dat…
Sentence