paper-with-me

Papers

From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

2025-05-20 · Yingli Shen, Wen Lai, Shuo Wang, Kangyang Luo, Alexander Fraser, Maosong Sun

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, multi-way parallel data, where identical content is aligned across multiple languages, provides stronger cross-lingual consistency and offers greater potential for improving multilingual performance. In this paper, we introduce a large-scale, high-quality multi-way parallel corpus, TED2025, based on TED Talks. The corpus spans 113 languages, with up to 50 languages aligned in parallel, ensuring extensive multilingual coverage. Using this dataset, we investigate best practices for leveraging multi-way parallel data to enhance LLMs, including strategies for continued pretraining, instruction tuning, and the analysis of key influencing factors. Experiments on six multilingual benchmarks show that models trained on multiway parallel data consistently outperform those trained on unaligned multilingual data.

📄 PDF Abstract BibTeX arXiv:2505.14045

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs

2025-08-07 · Franziska Weeber, Tanise Ceron, Sebastian Padó arxiv

Public opinion surveys show cross-cultural differences in political opinions between socio-cultural contexts. However, there is no clear evidence whether these differences translate to cross-lingual differences in multil…

Multilingual and Multimodal Topic Modelling with Pretrained Embeddings

2022-11-15 · COLING 2022 10 · Elaine Zosa, Lidia Pivovarova

This paper presents M3L-Contrast -- a novel multimodal multilingual (M3L) neural topic model for comparable data that maps texts from multiple languages and images into a shared topic space. Our model is trained jointly …

GlobalTrait: Personality Alignment of Multilingual Word Embeddings

2018-11-01 · Farhad Bin Siddique, Dario Bertero, Pascale Fung

We propose a multilingual model to recognize Big Five Personality traits from text data in four different languages: English, Spanish, Dutch and Italian. Our analysis shows that words having a similar semantic meaning in…

Multilingual Word EmbeddingsPersonality AlignmentWord Embeddings

Decoupled Alignment for Robust Plug-and-Play Adaptation

2024-06-03 · Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li 외

We introduce a low-resource safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning (SFT) or reinforcement learning from human feedback (RLHF). Our main idea is to …

Knowledge Distillation

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

2025-08-18 · Dayyán O'Brien, Bhavitvya Malik, Ona de Gibert, Pinzhen Chen 외 arxiv

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, …

Machine Translation