paper-with-me

홈 › Papers

Conditional Unigram Tokenization with Parallel Data

2025-07-10 · Gianluca Vico, Jindřinch Libovický arxiv

We introduce conditional unigram tokenization, a novel approach that extends unigram tokenization by conditioning target token probabilities on source-language tokens from parallel data. Given a fixed source tokenizer, our method learns a target tokenizer that maximizes cross-lingual semantic alignment. We evaluate our tokenizer on four language pairs across different families and resource levels, examining intrinsic properties and downstream performance on machine translation and language modeling. While our conditional tokenizer maintains comparable statistical properties to standard unigram tokenizers, results are mixed: we observe no improvements in machine translation quality, but find consistent perplexity reductions in language modeling. We hypothesize that quadratic scaling of conditional probability estimation with respect to the vocabulary size creates a data efficiency bottleneck. Our findings suggest that alternative parameterizations may be necessary for practical cross-lingual tokenization.

📄 PDF Abstract BibTeX arXiv:2507.07824

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Byte Pair Encoding is Suboptimal for Language Model Pretraining

2020-04-07 · Findings of the Association for Computational Linguistics 2020 · Kaj Bostrom, Greg Durrett

The success of pretrained transformer language models (LMs) in natural language processing has led to a wide range of pretraining setups. In particular, these models employ a variety of subword tokenization methods, most…

Language ModelingLanguage Modelling

Which Pieces Does Unigram Tokenization Really Need?

2025-12-14 · Sander Land, Yuval Pinter arxiv

The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegance, its implementation in practice is complex, limiting its adoption to …

Toward a Theory of Tokenization in LLMs

2024-04-12 · Nived Rajaraman, Jiantao Jiao, Kannan Ramchandran

While there has been a large body of research attempting to circumvent tokenization for language modeling (Clark et al., 2022; Xue et al., 2022), the current consensus is that it is a necessary initial step for designing…

Language ModelingLanguage Modelling

Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment

2025-08-11 · Saketh Reddy Vemula, Sandipan Dandapat, Dipti Misra Sharma, Parameswari Krishnamurthy arxiv

The relationship between tokenizer algorithm (e.g., Byte-Pair Encoding (BPE), Unigram), morphological alignment, tokenization quality (e.g., compression efficiency), and downstream performance remains largely unclear, pa…

Text Classification

Analyzing Cognitive Plausibility of Subword Tokenization

2023-10-20 · Lisa Beinborn, Yuval Pinter

Subword tokenization has become the de-facto standard for tokenization, although comparative evaluations of subword vocabulary quality across languages are scarce. Existing evaluation studies focus on the effect of a tok…