paper-with-me

홈 › Papers

Downstream Task-Oriented Neural Tokenizer Optimization with Vocabulary Restriction as Post Processing

2023-04-21 · Tatsuya Hiraoka, Tomoya Iwakura

This paper proposes a method to optimize tokenization for the performance improvement of already trained downstream models. Our method generates tokenization results attaining lower loss values of a given downstream model on the training data for restricting vocabularies and trains a tokenizer reproducing the tokenization results. Therefore, our method can be applied to variety of tokenization methods, while existing work cannot due to the simultaneous learning of the tokenizer and the downstream model. This paper proposes an example of the BiLSTM-based tokenizer with vocabulary restriction, which can capture wider contextual information for the tokenization process than non-neural-based tokenization methods used in existing work. Experimental results on text classification in Japanese, Chinese, and English text classification tasks show that the proposed method improves performance compared to the existing methods for tokenization optimization.

📄 PDF Abstract BibTeX arXiv:2304.10808

Code (0)

등록된 구현이 없습니다.

Tasks

text-classificationText Classification

Similar Papers 제목 키워드 기반

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

2026-09-08 · Siting Li, Zhengyang Wang, Simon Shaolei Du, Xi Chen 외 hf

Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how vi…

Continual Pretraining

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment

2026-06-25 · Sander Land arxiv

The Unigram tokenizer uses an elegant representation which makes it straightforward to edit vocabularies, but its training is comparatively heavy and complex. We introduce MinGram (Minimalist Unigram), which keeps the to…

A Vocabulary-Free Multilingual Neural Tokenizer for End-to-End Task Learning

2022-04-22 · RepL4NLP (ACL) 2022 5 · Md Mofijul Islam, Gustavo Aguilar, Pragaash Ponnusamy, Clint Solomon Mathialagan 외

Subword tokenization is a commonly used input pre-processing step in most recent NLP models. However, it limits the models' ability to leverage end-to-end task learning. Its frequency-based vocabulary creation compromise…

DiversitySentiment Analysis

Tokenizer Choice For LLM Training: Negligible or Crucial?

2023-10-12 · Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann 외

The recent success of Large Language Models (LLMs) has been predominantly driven by curating the training dataset composition, scaling of model architectures and dataset sizes and advancements in pretraining objectives, …

Tokenization Impacts Multilingual Language Modeling: Assessing Vocabulary Allocation and Overlap Across Languages

2023-05-26 · Tomasz Limisiewicz, Jiří Balhar, David Mareček

Multilingual language models have recently gained attention as a promising solution for representing multiple languages in a single model. In this paper, we propose new criteria to evaluate the quality of lexical represe…

Language ModelingLanguage ModellingNERPOS+2