paper-with-me

홈 › Papers

From English-Centric to Effective Bilingual: LLMs with Custom Tokenizers for Underrepresented Languages

2024-10-24 · Artur Kiulian, Anton Polishko, Mykola Khandoga, Yevhen Kostiuk, Guillermo Gabrielli, Łukasz Gagała, Fadi Zaraket, Qusai Abu Obaida, Hrishikesh Garud, Wendy Wing Yee Mak, Dmytro Chaplynskyi, Selma Belhadj Amor, Grigol Peradze

In this paper, we propose a model-agnostic cost-effective approach to developing bilingual base large language models (LLMs) to support English and any target language. The method includes vocabulary expansion, initialization of new embeddings, model training and evaluation. We performed our experiments with three languages, each using a non-Latin script - Ukrainian, Arabic, and Georgian. Our approach demonstrates improved language performance while reducing computational costs. It mitigates the disproportionate penalization of underrepresented languages, promoting fairness and minimizing adverse phenomena such as code-switching and broken grammar. Additionally, we introduce new metrics to evaluate language quality, revealing that vocabulary size significantly impacts the quality of generated text.

📄 PDF Abstract BibTeX arXiv:2410.18836

Code (0)

등록된 구현이 없습니다.

Tasks

Fairness

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

CroissantLLM: A Truly Bilingual French-English Language Model

2024-02-01 · Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison 외

We introduce CroissantLLM, a 1.3B language model pretrained on a set of 3T English and French tokens, to bring to the research and industrial community a high-performance, fully open-sourced bilingual model that runs swi…

Language ModelingLanguage ModellingLarge Language Modelmodel

Improving Bilingual Capabilities of Language Models to Support Diverse Linguistic Practices in Education

2024-11-06 · Anand Syamkumar, Nora Tseng, Kaycie Barron, Shanglin Yang 외

Large language models (LLMs) offer promise in generating educational content, providing instructor feedback, and reducing teacher workload on assessments. While prior studies have focused on studying LLM-powered learning…

Dólares or Dollars? Unraveling the Bilingual Prowess of Financial LLMs Between Spanish and English

2024-02-12 · Xiao Zhang, Ruoyu Xiang, Chenhan Yuan, Duanyu Feng 외

Despite Spanish's pivotal role in the global finance industry, a pronounced gap exists in Spanish financial natural language processing (NLP) and application studies compared to English, especially in the era of large la…

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

2026-08-18 · Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov 외 arxiv

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks e…

Information Extraction

Findings of the WMT 2020 Shared Task on Chat Translation

2020-11-01 · WMT (EMNLP) 2020 11 · M. Amin Farajian, António V. Lopes, André F. T. Martins, Sameen Maruf 외

We report the results of the first edition of the WMT shared task on chat translation. The task consisted of translating bilingual conversational text, in particular customer support chats for the English-German language…

de-enTranslation