paper-with-me

홈 › Papers

Gold Panning in Vocabulary: An Adaptive Method for Vocabulary Expansion of Domain-Specific LLMs

2024-10-02 · Chengyuan Liu, Shihang Wang, Lizhi Qing, Kun Kuang, Yangyang Kang, Changlong Sun, Fei Wu

While Large Language Models (LLMs) demonstrate impressive generation abilities, they frequently struggle when it comes to specialized domains due to their limited domain-specific knowledge. Studies on domain-specific LLMs resort to expanding the vocabulary before fine-tuning on domain-specific corpus, aiming to decrease the sequence length and enhance efficiency during decoding, without thoroughly investigating the results of vocabulary expansion to LLMs over different domains. Our pilot study reveals that expansion with only a subset of the entire vocabulary may lead to superior performance. Guided by the discovery, this paper explores how to identify a vocabulary subset to achieve the optimal results. We introduce VEGAD, an adaptive method that automatically identifies valuable words from a given domain vocabulary. Our method has been validated through experiments on three Chinese datasets, demonstrating its effectiveness. Additionally, we have undertaken comprehensive analyses of the method. The selection of a optimal subset for expansion has shown to enhance performance on both domain-specific tasks and general tasks, showcasing the potential of VEGAD.

📄 PDF Abstract BibTeX arXiv:2410.01188

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Can We Effectively Expand the Vocabulary of LLMs with 0.01GB of Target Language Text?

2024-06-17 · Atsuki Yamaguchi, Aline Villavicencio, Nikolaos Aletras

Large language models (LLMs) have shown remarkable capabilities in many languages beyond English. Yet, LLMs require more inference steps when generating non-English text due to their reliance on English-centric tokenizer…

Cross-Lingual Transfer

Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion

2024-12-16 · Jianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang 외

This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or C…

Scaling LLM Pre-training with Vocabulary Curriculum

2025-02-25 · Fangyuan Yu

Modern language models rely on static vocabularies, fixed before pretraining, in contrast to the adaptive vocabulary acquisition observed in human language learning. To bridge this gap, we introduce vocabulary curriculum…

Model Optimization

Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

2026-04-17 · Maitrey Mehta, Nishant Subramani, Zhichao Xu, Ashim Gupta 외 arxiv

All languages are equal; when it comes to tokenization, some are more equal than others. Tokens are the hidden currency that dictate the cost and latency of access to contemporary LLMs. However, many languages written in…

Now It Sounds Like You: Learning Personalized Vocabulary On Device

2023-05-05 · Sid Wang, Ashish Shenoy, Pierce Chuang, John Nguyen

In recent years, Federated Learning (FL) has shown significant advancements in its ability to perform various natural language processing (NLP) tasks. This work focuses on applying personalized FL for on-device language …

Federated LearningLanguage ModelingLanguage Modelling