paper-with-me

홈 › Papers

False Friends Are Not Foes: Investigating Vocabulary Overlap in Multilingual Language Models

2025-09-23 · Julie Kallini, Dan Jurafsky, Christopher Potts, Martijn Bartelds arxiv

Subword tokenizers trained on multilingual corpora naturally produce overlapping tokens across languages. Does token overlap facilitate cross-lingual transfer or instead introduce interference between languages? Prior work offers mixed evidence, partly due to varied setups and confounders, such as token frequency or subword segmentation granularity. To address this question, we devise a controlled experiment where we train bilingual autoregressive models on multiple language pairs under systematically varied vocabulary overlap settings. Crucially, we explore a new dimension to understanding how overlap affects transfer: the semantic similarity of tokens shared across languages. We first analyze our models' hidden representations and find that overlap of any kind creates embedding spaces that capture cross-lingual semantic relationships, while this effect is much weaker in models with disjoint vocabularies. On XNLI and XQuAD, we find that models with overlap outperform models with disjoint vocabularies, and that transfer performance generally improves as overlap increases. Overall, our findings highlight the advantages of token overlap in multilingual models and show that substantial shared vocabulary remains a beneficial design choice for multilingual tokenizers.

📄 PDF Abstract BibTeX arXiv:2509.18750

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferSemantic Similarity

Similar Papers 제목 키워드 기반

Is Cross-Lingual Transfer in Bilingual Models Human-Like? A Study with Overlapping Word Forms in Dutch and English

2026-04-08 · Iza Škrjanec, Irene Elisabeth Winther, Vera Demberg, Stefan L. Frank arxiv

Bilingual speakers show cross-lingual activation during reading, especially for words with shared surface form. Cognates (friends) typically lead to facilitation, whereas interlingual homographs (false friends) cause int…

Cross-Lingual Transfer

Friends and Foes in Learning from Noisy Labels

2021-03-28 · Yifan Zhou, Yifan Ge, Jianxin Wu

Learning from examples with noisy labels has attracted increasing attention recently. But, this paper will show that the commonly used CIFAR-based datasets and the accuracy evaluation metric used in the literature are bo…

Self-Supervised Learningvalid

Leveraging Trust and Distrust in Recommender Systems via Deep Learning

2019-05-31 · Dimitrios Rafailidis

The data scarcity of user preferences and the cold-start problem often appear in real-world applications and limit the recommendation accuracy of collaborative filtering strategies. Leveraging the selections of social fr…

Collaborative FilteringDeep LearningRecommendation Systems

Modeling Friends and Foes

2018-06-30 · Pedro A. Ortega, Shane Legg

How can one detect friendly and adversarial behavior from raw data? Detecting whether an environment is a friend, a foe, or anything in between, remains a poorly understood yet desirable ability for safe and robust agent…

Automatically Building a Multilingual Lexicon of False Friends With No Supervision

2020-05-01 · LREC 2020 5 · Ana Sabina Uban, Liviu P. Dinu

Cognate words, defined as words in different languages which derive from a common etymon, can be useful for language learners, who can leverage the orthographical similarity of cognates to more easily understand a text i…

Cross-Lingual Word EmbeddingsLanguage AcquisitionWord Embeddings