paper-with-me

홈 › Papers

Toucan: Token-Aware Character Level Language Modeling

2023-11-15 · William Fleshman, Benjamin Van Durme

Character-level language models obviate the need for separately trained tokenizers, but efficiency suffers from longer sequence lengths. Learning to combine character representations into tokens has made training these models more efficient, but they still require decoding characters individually. We propose Toucan, an augmentation to character-level models to make them "token-aware". Comparing our method to prior work, we demonstrate significant speed-ups in character generation without a loss in language modeling performance. We then explore differences between our learned dynamic tokenization of character sequences with popular fixed vocabulary solutions such as Byte-Pair Encoding and WordPiece, finding our approach leads to a greater amount of longer sequences tokenized as single items. Our project and code are available at https://nlp.jhu.edu/nuggets/.

📄 PDF Abstract BibTeX arXiv:2311.08620

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Toucan: Many-to-Many Translation for 150 African Language Pairs

2024-07-05 · AbdelRahim Elmadany, Ife Adebara, Muhammad Abdul-Mageed

We address a notable gap in Natural Language Processing (NLP) by introducing a collection of resources designed to improve Machine Translation (MT) for low-resource languages, with a specific focus on African languages. …

Machine TranslationTranslation

TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments

2025-10-01 · Zhangchen Xu, Adriana Meza Soria, Shawn Tan, Anurag Roy 외 arxiv

Large Language Model (LLM) agents are rapidly emerging as powerful systems for automating tasks across domains. Yet progress in the open-source community is constrained by the lack of high quality permissively licensed t…

Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning

2024-11-26 · Zhu Xu, Zhiqiang Zhao, Zihan Zhang, Yuchi Liu 외

Tokenization methods like Byte-Pair Encoding (BPE) enhance computational efficiency in large language models (LLMs) but often obscure internal character structures within tokens. This limitation hinders LLMs' ability to …

Computational EfficiencyPositionPredictionSpelling Correction

Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay

2026-02-06 · Duygu Altinok arxiv

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fideli…

Dependency ParsingSentiment Analysis

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models

2025-08-07 · Chenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu 외 arxiv

While large language models (LLMs) have demonstrated remarkable performance on high-level semantic tasks, they often struggle with fine-grained, token-level understanding and structural reasoning--capabilities that are e…

Synthetic Data Generation