paper-with-me

홈 › Papers

Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness

2026-03-11 · Zhipeng Yang, Shu Yang, Lijie Hu, Di Wang arxiv

Large language models (LLMs) trained with canonical tokenization exhibit surprising robustness to non-canonical inputs such as character-level tokenization, yet the mechanisms underlying this robustness remain unclear. We study this phenomenon through mechanistic interpretability and identify a core process we term word recovery. We first introduce a decoding-based method to detect word recovery, showing that hidden states reconstruct canonical word-level token identities from character-level inputs. We then provide causal evidence by removing the corresponding subspace from hidden states, which consistently degrades downstream task performance. Finally, we conduct a fine-grained attention analysis and show that in-group attention among characters belonging to the same canonical token is critical for word recovery: masking such attention in early layers substantially reduces both recovery scores and task performance. Together, our findings provide a mechanistic explanation for tokenization robustness and identify word recovery as a key mechanism enabling LLMs to process character-level inputs.

📄 PDF Abstract BibTeX arXiv:2603.10771

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Chinese Word Boundary Recovery through Character Alignment Projection

2026-05-27 · Lusha Wang, Yuchen Li, Su Yuan, Jungyeul Park arxiv

Chinese word segmentation is especially fragile in non-standard text, where language learner errors and other character-level divergences disrupt the word boundaries assumed by downstream annotation and evaluation. This …

Chinese Word Segmentation

Syllable Subword Tokens for Open Vocabulary Speech Recognition in Malayalam

2023-01-17 · Kavya Manohar, A. R. Jayan, Rajeev Rajan

In a hybrid automatic speech recognition (ASR) system, a pronunciation lexicon (PL) and a language model (LM) are essential to correctly retrieve spoken word sequences. Being a morphologically complex language, the vocab…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Chain Association-based Attacking and Shielding Natural Language Processing Systems

2024-11-12 · Jiacheng Huang, Long Chen

Association as a gift enables people do not have to mention something in completely straightforward words and allows others to understand what they intend to refer to. In this paper, we propose a chain association-based …

Adversarial Attack

A Joint Model for Word Embedding and Word Morphology

2016-06-08 · WS 2016 8 · Kris Cao, Marek Rei

This paper presents a joint model for performing unsupervised morphological analysis on words, and learning a character-level composition function from morphemes to word embeddings. Our model splits individual words into…

Morphological AnalysisWord Embeddings

Context-based out-of-vocabulary word recovery for ASR systems in Indian languages

2022-06-09 · Arun Baby, Saranya Vinnaitherthan, Akhil Kerhalkar, Pranav Jawale 외

Detecting and recovering out-of-vocabulary (OOV) words is always challenging for Automatic Speech Recognition (ASR) systems. Many existing methods focus on modeling OOV words by modifying acoustic and language models and…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3