paper-with-me

홈 › Papers

Adversarial Tokenization

2025-03-04 · Renato Lui Geh, Zilei Shao, Guy Van Den Broeck

Current LLM pipelines account for only one possible tokenization for a given string, ignoring exponentially many alternative tokenizations during training and inference. For example, the standard Llama3 tokenization of penguin is [p,enguin], yet [peng,uin] is another perfectly valid alternative. In this paper, we show that despite LLMs being trained solely on one tokenization, they still retain semantic understanding of other tokenizations, raising questions about their implications in LLM safety. Put succinctly, we answer the following question: can we adversarially tokenize an obviously malicious string to evade safety and alignment restrictions? We show that not only is adversarial tokenization an effective yet previously neglected axis of attack, but it is also competitive against existing state-of-the-art adversarial approaches without changing the text of the harmful request. We empirically validate this exploit across three state-of-the-art LLMs and adversarial datasets, revealing a previously unknown vulnerability in subword models.

📄 PDF Abstract BibTeX arXiv:2503.02174

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

Adversarially-Refined VQ-GAN with Dense Motion Tokenization for Spatio-Temporal Heatmaps

2025-09-23 · Gabriel Maldonado, Narges Rashvand, Armin Danesh Pazho, Ghazal Alinezhad Noghre 외 arxiv

Continuous human motion understanding remains a core challenge in computer vision due to its high dimensionality and inherent redundancy. Efficient compression and representation are crucial for analyzing complex motion …

Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization

2024-05-27 · Dixuan Wang, Yanda Li, Junyuan Jiang, Zepeng Ding 외

Large Language Models (LLMs) have shown remarkable capabilities in language understanding and generation. Nonetheless, it was also witnessed that LLMs tend to produce inaccurate responses to specific queries. This defici…

Extend Adversarial Policy Against Neural Machine Translation via Unknown Token

2025-01-21 · Wei Zou, ShuJian Huang, Jiajun Chen

Generating adversarial examples contributes to mainstream neural machine translation~(NMT) robustness. However, popular adversarial policies are apt for fixed tokenization, hindering its efficacy for common character per…

Machine TranslationNMTReinforcement Learning (RL)Translation

Neural Sign Language Translation by Learning Tokenization

2020-02-02 · Alptekin Orbay, Lale Akarun

Sign Language Translation has attained considerable success recently, raising hopes for improved communication with the Deaf. A pre-processing step called tokenization improves the success of translations. Tokens can be …

Sign Language TranslationTransfer LearningTranslation

End-to-End Training for Unified Tokenization and Latent Denoising

2026-03-23 · Shivam Duggal, Xingjian Bai, Zongze Wu, Richard Zhang 외 arxiv

Latent diffusion models (LDMs) enable high-fidelity synthesis by operating in learned latent spaces. However, training state-of-the-art LDMs requires complex staging: a tokenizer must be trained first, before the diffusi…