paper-with-me

홈 › Papers

Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking

2023-11-16 · Nan Xu, Fei Wang, Ben Zhou, Bang Zheng Li, Chaowei Xiao, Muhao Chen

While large language models (LLMs) have demonstrated increasing power, they have also given rise to a wide range of harmful behaviors. As representatives, jailbreak attacks can provoke harmful or unethical responses from LLMs, even after safety alignment. In this paper, we investigate a novel category of jailbreak attacks specifically designed to target the cognitive structure and processes of LLMs. Specifically, we analyze the safety vulnerability of LLMs in the face of (1) multilingual cognitive overload, (2) veiled expression, and (3) effect-to-cause reasoning. Different from previous jailbreak attacks, our proposed cognitive overload is a black-box attack with no need for knowledge of model architecture or access to model weights. Experiments conducted on AdvBench and MasterKey reveal that various LLMs, including both popular open-source model Llama 2 and the proprietary model ChatGPT, can be compromised through cognitive overload. Motivated by cognitive psychology work on managing cognitive load, we further investigate defending cognitive overload attack from two perspectives. Empirical studies show that our cognitive overload from three perspectives can jailbreak all studied LLMs successfully, while existing defense strategies can hardly mitigate the caused malicious uses effectively.

📄 PDF Abstract BibTeX arXiv:2311.09827

Code (0)

등록된 구현이 없습니다.

Tasks

Safety Alignment

Similar Papers 제목 키워드 기반

Deep Learning Overloaded Vehicle Identification for Long Span Bridges Based on Structural Health Monitoring Data

2023-09-04 · Yuqin Li, Jun Liu, Shengliang Zhong, Licheng Zhou 외

Overloaded vehicles bring great harm to transportation infrastructures. BWIM (bridge weigh-in-motion) method for overloaded vehicle identification is getting more popular because it can be implemented without interruptio…

Deep LearningStructural Health Monitoring

Deep Learning-Aided Projected Gradient Detector for Massive Overloaded MIMO Channels

2018-06-28 · Satoshi Takabe, Masayuki Imanishi, Tadashi Wadayama, Kazunori Hayashi

The paper presents a deep learning-aided iterative detection algorithm for massive overloaded MIMO systems. Since the proposed algorithm is based on the projected gradient descent method with trainable parameters, it is …

Deep Learning

InfoFlood: Jailbreaking Large Language Models with Information Overload

2025-06-13 · Advait Yadav, Haibo Jin, Man Luo, Jun Zhuang 외

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains. However, their potential to generate harmful responses has raised significant societal and regulatory concerns, especially wh…

Trainable Projected Gradient Detector for Massive Overloaded MIMO Channels: Data-driven Tuning Approach

2018-12-25 · Satoshi Takabe, Masayuki Imanishi, Tadashi Wadayama, Ryo Hayakawa 외

This paper presents a deep learning-aided iterative detection algorithm for massive overloaded multiple-input multiple-output (MIMO) systems where the number of transmit antennas $n$ is larger than that of receive antenn…

parameter estimation

Modeling Communication to Coordinate Perspectives in Cooperation

2021-06-03 · Stephanie Stacy, Chenfei Li, Minglu Zhao, Yiling Yun 외

Communication is highly overloaded. Despite this, even young children are good at leveraging context to understand ambiguous signals. We propose a computational account of overloaded signaling from a shared agency perspe…