paper-with-me

Papers

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning

2025-08-05 · Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang arxiv

Defending large language models (LLMs) against jailbreak attacks is essential for their safe and reliable deployment. Existing defenses often rely on shallow pattern matching, which struggles to generalize to novel and unseen attack strategies. To address this challenge, we propose the Cognitive-Driven Defense (CDD) framework, which targets the underlying structure of jailbreak prompts by applying meta-operations, defined as basic manipulations that conceal harmful intent.CDD emulates human cognitive reasoning through a structured reasoning chain. It begins with a global perception of the prompt and follows with a localized analysis to uncover hidden manipulations. By applying supervised fine-tuning on this structured chain, the model learns to identify and reason about known manipulation patterns. To enhance generalization to unseen threats, an entropy-guided reinforcement learning algorithm (EG-GRPO) is introduced to encourage exploration of new types and variants of meta-operations. Experiments demonstrate that CDD can achieve state-of-the-art defense performance and exhibit strong generalization to unseen jailbreak attacks.

📄 PDF Abstract BibTeX arXiv:2508.03054

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Memory Beyond Recall: A Dual-Process Cognitive Memory System for Self-Evolving LLM Agents

2026-06-08 · Tianxiang Fei, Mingyang Song, Mao Zheng, Xiang Yu arxiv

Long-term memory for an LLM agent is more than retrieving the right passage at the right time. Current memory systems collapse belief revision, causal coupling, and cross-domain abstraction into a single retrieval surfac…

Narrative-Centered Emotional Reflection: Scaffolding Autonomous Emotional Literacy with AI

2025-04-29 · Shou-Tzu Han

Reflexion is an AI-powered platform designed to enable structured emotional self-reflection at scale. By integrating real-time emotion detection, layered reflective prompting, and metaphorical storytelling generation, Re…

Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety

2025-10-11 · Yuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao 외 arxiv

As large language models (LLMs) are increasingly deployed for complex reasoning tasks, Long Chain-of-Thought (Long-CoT) prompting has emerged as a key paradigm for structured inference. Despite early-stage safeguards ena…

Can Large Language Models Simulate Human Cognition Beyond Behavioral Imitation?

2026-03-29 · Yuxuan Gu, Lunjun Liu, Xiaocheng Feng, Kun Zhu 외 arxiv

An essential problem in artificial intelligence is whether LLMs can simulate human cognition or merely imitate surface-level behaviors, while existing datasets suffer from either synthetic reasoning traces or population-…

Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech

2025-12-15 · Dylan Phelps, Rodrigo Wilkens, Edward Gow-Smith, Lilian Hubner 외 arxiv

Alzheimer's Disease (AD) is a progressive neurodegenerative condition that adversely affects cognitive abilities. Language-related changes can be automatically identified through the analysis of outputs from linguistic a…

Alzheimer's Disease DetectionSemantic Similarity