paper-with-me

홈 › Papers

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

2025-08-14 · Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva arxiv

Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to study the causes and training dynamics of this phenomenon by training small Transformer models from scratch on synthetic multilingual datasets. Depending on (1) the correlation between facts and the language they were learned in (informativeness), and (2) the ease of language identification (extractability), models either develop unified representations across languages or separate representations; only when representations are unified do facts transfer across languages. Based on these insights, we propose a unifying perspective which explains a range of prior observations concerning cross-lingual transfer in multilingual LLMs. Our work shows controlled settings can shed light on pre-training dynamics and suggests methods to encourage representational unification as part of training that would improve LLMs' cross-lingual transfer.

📄 PDF Abstract BibTeX arXiv:2508.11017

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

From Rosetta to Match-Up: A Paired Corpus of Linguistic Puzzles with Human and LLM Benchmarks

2026-05-13 · Neh Majmudar, Anne Huang, Jinfan Frank Hu, Elena Filatova arxiv

In this paper, we examine linguistic puzzles used in high school linguistics competitions, focusing on two common formats: Rosetta Stone and Match-Up. We propose a systematic procedure for converting existing Rosetta Sto…

Rosetta Stone Linguistic Problems

2013-08-01 · WS 2013 8 · Bozhidar Bozhanov, Ivan Derzhanski

Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning

2025-01-01 · CVPR 2025 1 · Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

Egocentric and exocentric perspectives of human action differ significantly, yet overcoming this extreme viewpoint gap is critical for applications in augmented reality and robotics. We propose ViewpointRosetta, an a…

Action RecognitionContrastive LearningRepresentation Learning

Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning

2025-03-25 · Shaun Baek, Shaun Esua-Mensah, Cyrus Tsui, Sejan Vigneswaralingam 외

Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resource settings and in tasks requiring deep logical reasoning. This research introduces Rosetta…

Language ModelingLanguage ModellingLarge Language ModelLogical Reasoning+1

A Rosetta Stone for AI Benchmarks

2025-11-28 · Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie 외 arxiv

Most AI benchmarks saturate within years or even months after they are introduced, making it hard to study long-run trends in AI capabilities. To address this challenge, we build a statistical framework that stitches ben…