paper-with-me

홈 › Papers

To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models

2025-02-16 · Zihao Zhu, Hongbao Zhang, Ruotong Wang, Ke Xu, Siwei Lyu, Baoyuan Wu

Large Reasoning Models (LRMs) are designed to solve complex tasks by generating explicit reasoning traces before producing final answers. However, we reveal a critical vulnerability in LRMs -- termed Unthinking Vulnerability -- wherein the thinking process can be bypassed by manipulating special delimiter tokens. It is empirically demonstrated to be widespread across mainstream LRMs, posing both a significant risk and potential utility, depending on how it is exploited. In this paper, we systematically investigate this vulnerability from both malicious and beneficial perspectives. On the malicious side, we introduce Breaking of Thought (BoT), a novel attack that enables adversaries to bypass the thinking process of LRMs, thereby compromising their reliability and availability. We present two variants of BoT: a training-based version that injects backdoor during the fine-tuning stage, and a training-free version based on adversarial attack during the inference stage. As a potential defense, we propose thinking recovery alignment to partially mitigate the vulnerability. On the beneficial side, we introduce Monitoring of Thought (MoT), a plug-and-play framework that allows model owners to enhance efficiency and safety. It is implemented by leveraging the same vulnerability to dynamically terminate redundant or risky reasoning through external monitoring. Extensive experiments show that BoT poses a significant threat to reasoning reliability, while MoT provides a practical solution for preventing overthinking and jailbreaking. Our findings expose an inherent flaw in current LRM architectures and underscore the need for more robust reasoning systems in the future.

📄 PDF Abstract BibTeX arXiv:2502.12202

Code (2)

zihao-ai/bot 공식 구현 pytorch
zihao-ai/unthinking_vulnerability 공식 구현 pytorch

Tasks

Adversarial AttackBackdoor Attack

Similar Papers 제목 키워드 기반

Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion

2025-09-25 · Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju 외 arxiv

Understanding persuasion is critical for the safety and reliability of multi-agent systems built on large language models (LLMs). This paper studies persuasion dynamics by contrasting general LLMs with Large Reasoning Mo…

Think Fast, Think Slow, Think Critical: Designing an Automated Propaganda Detection Tool

2024-02-29 · Liudmila Zavolokina, Kilian Sprenkamp, Zoya Katashinskaya, Daniel Gordon Jones 외

In today's digital age, characterized by rapid news consumption and increasing vulnerability to propaganda, fostering citizens' critical thinking is crucial for stable democracies. This paper introduces the design of Cla…

ArticlesPropaganda detection

How critically can an AI think? A framework for evaluating the quality of thinking of generative artificial intelligence

2024-06-20 · Luke Zaphir, Jason M. Lodge, Jacinta Lisec, Dom McGrath 외

Generative AI such as those with large language models have created opportunities for innovative assessment design practices. Due to recent technological developments, there is a need to know the limits and capabilities …

How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs

2026-01-07 · Su-Hyeon Kim, Hyundong Jin, Yejin Lee, Yo-Sub Han arxiv

Large Reasoning Models (LRMs) achieve remarkable success through explicit thinking steps, yet the thinking steps introduce a novel risk by potentially amplifying unsafe behaviors. Despite this vulnerability, conventional…

Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning

2025-09-17 · Haodong Zhao, Chenyan Zhao, Yansi Li, Zhuosheng Zhang 외 arxiv

The capacity of Large Language Models (LLMs) to reason is fundamental to their application in complex, knowledge-intensive domains. In real-world scenarios, LLMs are often augmented with external information that can be …