paper-with-me

홈 › Papers

CharBench: Evaluating the Role of Tokenization in Character-Level Tasks

2025-08-04 · Omri Uzan, Yuval Pinter arxiv

Tasks that require character-level reasoning, such as counting or locating characters within words, remain challenging for contemporary language models. A common conjecture is that language models' reliance on subword units, rather than characters, contributes to their struggles with character-level tasks, yet recent studies offer conflicting conclusions about the role of tokenization, leaving its impact unclear. To address this gap, we introduce CharBench, a comprehensive benchmark of character-level tasks that is two orders of magnitude larger than existing alternatives. We evaluate a diverse range of leading open-weight and proprietary models on CharBench and find that it presents a significant challenge to modern LLMs, with an average accuracy of 43.6% and 32.3% on some tasks. We present an in-depth analysis of how intrinsic properties of words and their segmentations into tokens correspond to model performance. For counting tasks, we find that tokenization properties are weakly correlated with correctness, while the length of the queried word and the actual character count play a more significant part. In contrast, for tasks requiring intra-word positional understanding, performance is negatively correlated with the length of the token containing the queried character, suggesting that longer tokens obscure character position information for LLMs. We encourage future work to build on the benchmark and evaluation methodology introduced here as tools for improving model performance on such tasks.

📄 PDF Abstract BibTeX arXiv:2508.02591

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tokenization Strategies for Low-Resource Agglutinative Languages in Word2Vec: Case Study on Turkish and Finnish

2025-08-27 · Jinfan Frank Hu arxiv

Tokenization plays a critical role in processing agglutinative languages, where a single word can encode multiple morphemes carrying syntactic and semantic information. This study evaluates the impact of various tokeniza…

How Important Is Tokenization in French Medical Masked Language Models?

2024-02-22 · Yanis Labrak, Adrien Bazoge, Beatrice Daille, Mickael Rouvier 외

Subword tokenization has become the prevailing standard in the field of natural language processing (NLP) over recent years, primarily due to the widespread utilization of pre-trained language models. This shift began wi…

Evaluating Interval-based Tokenization for Pitch Representation in Symbolic Music Analysis

2025-01-08 · Dinh-Viet-Toan Le, Louis Bigo, Mikaela Keller

Symbolic music analysis tasks are often performed by models originally developed for Natural Language Processing, such as Transformers. Such models require the input data to be represented as sequences, which is achieved…

A Study on Dialog Act Recognition using Character-Level Tokenization

2018-05-18 · Eugénio Ribeiro, Ricardo Ribeiro, David Martins de Matos

Dialog act recognition is an important step for dialog systems since it reveals the intention behind the uttered words. Most approaches on the task use word-level tokenization. In contrast, this paper explores the use of…

How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis

2025-11-05 · Ahmed Mostafa, Raisul Arefin Nahid, Samuel Mulder arxiv

Tokenization is fundamental in assembly code analysis, impacting intrinsic characteristics like vocabulary size, semantic coverage, and extrinsic performance in downstream tasks. Despite its significance, tokenization in…