paper-with-me

홈 › Papers

How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?

2024-06-17 · Atsuki Yamaguchi, Aline Villavicencio, Nikolaos Aletras

Large language models (LLMs) have shown remarkable capabilities in many languages beyond English. Yet, LLMs require more inference steps when generating non-English text due to their reliance on English-centric tokenizers and vocabulary, resulting in higher usage costs to non-English speakers. Vocabulary expansion with target language tokens is a widely used cross-lingual vocabulary adaptation approach to remedy this issue. Despite its effectiveness in inference speedup, previous work on vocabulary expansion has focused on high-resource settings assuming access to a substantial amount of target language data to effectively initialize the embeddings of the new tokens and adapt the LLM to the target language. However, vocabulary expansion in low-resource settings has yet to be explored. In this paper, we investigate vocabulary expansion in low-resource settings by considering embedding initialization methods and continual pre-training strategies. Through extensive experiments across typologically diverse languages, tasks and models, we establish a set of strategies to perform vocabulary expansion for faster inference, maintaining competitive downstream performance to baselines with only 30K sentences ($\sim$0.01GB text data) from the target language.

📄 PDF Abstract BibTeX arXiv:2406.11477

Code (2)

gucci-j/lowres-cva 공식 구현 pytorch
gucci-j/lowres-cve 공식 구현 pytorch

Tasks

Cross-Lingual Transfer

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Effective vocabulary expansion of multilingual language models for extremely low-resource languages

2026-02-10 · Jianyu Zheng arxiv

Multilingual pre-trained language models(mPLMs) offer significant benefits for many low-resource languages. To further expand the range of languages these models can support, many works focus on continued pre-training of…

POS Tagging

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

2024-11-14 · Zhengyi Wang, Jonathan Lorraine, Yikai Wang, Hang Su 외

This work explores expanding the capabilities of large language models (LLMs) pretrained on text to generate 3D meshes within a unified model. This offers key advantages of (1) leveraging spatial knowledge already embedd…

3D GenerationText Generation

CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems

2025-06-24 · Haochen Zhang, Tianyi Zhang, Junze Yin, Oren Gal 외

Recommender systems play a pivotal role in providing relevant content to users. With the rapid development of large language models (LLMs), researchers have begun utilizing LLMs to build more powerful recommender systems…

Recommendation Systems

Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models

2023-10-23 · Gonzalo Martínez, Javier Conde, Elena Merino-Gómez, Beatriz Bermúdez-Margaretto 외

Vocabulary tests, once a cornerstone of language modeling evaluation, have been largely overlooked in the current landscape of Large Language Models (LLMs) like Llama, Mistral, and GPT. While most LLM evaluation benchmar…

Language ModelingLanguage Modelling

Learning to Link Grammar and Encyclopedic Information of Assist ESL Learners

2019-07-01 · ACL 2019 7 · Jhih-Jie Chen, Ching-Yu Yang, Peichen Ho, Ming Chiao Tsai 외

We introduce a system aimed at improving and expanding second language learners{'} English vocabulary. In addition to word definitions, we provide rich lexical information such as collocations and grammar patterns for ta…