paper-with-me

홈 › Papers

Dagger Behind Smile: Fool LLMs with a Happy Ending Story

2025-01-19 · Xurui Song, Zhixin Xie, Shuo Huai, Jiayi Kong, Jun Luo

The wide adoption of Large Language Models (LLMs) has attracted significant attention from $\textit{jailbreak}$ attacks, where adversarial prompts crafted through optimization or manual design exploit LLMs to generate malicious contents. However, optimization-based attacks have limited efficiency and transferability, while existing manual designs are either easily detectable or demand intricate interactions with LLMs. In this paper, we first point out a novel perspective for jailbreak attacks: LLMs are more responsive to $\textit{positive}$ prompts. Based on this, we deploy Happy Ending Attack (HEA) to wrap up a malicious request in a scenario template involving a positive prompt formed mainly via a $\textit{happy ending}$, it thus fools LLMs into jailbreaking either immediately or at a follow-up malicious request.This has made HEA both efficient and effective, as it requires only up to two turns to fully jailbreak LLMs. Extensive experiments show that our HEA can successfully jailbreak on state-of-the-art LLMs, including GPT-4o, Llama3-70b, Gemini-pro, and achieves 88.79\% attack success rate on average. We also provide quantitative explanations for the success of HEA.

📄 PDF Abstract BibTeX arXiv:2501.13115

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Spontaneous vs. Posed smiles - can we tell the difference?

2016-05-23 · Bappaditya Mandal, Nizar Ouarti

Smile is an irrefutable expression that shows the physical state of the mind in both true and deceptive ways. Generally, it shows happy state of the mind, however, `smiles' can be deceptive, for example people can give a…

Optical Flow Estimation

SMILe: Scalable Meta Inverse Reinforcement Learning through Context-Conditional Policies

2019-12-01 · NeurIPS 2019 12 · Seyed Kamyar Seyed Ghasemipour, Shixiang (Shane) Gu, Richard Zemel

Imitation Learning (IL) has been successfully applied to complex sequential decision-making problems where standard Reinforcement Learning (RL) algorithms fail. A number of recent methods extend IL to few-shot learning s…

continuous-controlContinuous ControlDecision MakingFew-Shot Learning+5

SMILE: Multimodal Dataset for Understanding Laughter in Video with Language Models

2023-12-15 · Lee Hyun, Kim Sung-Bin, Seungju Han, Youngjae Yu 외

Despite the recent advances of the artificial intelligence, building social intelligence remains a challenge. Among social signals, laughter is one of the distinctive expressions that occurs during social interactions be…

Video Understanding

How to Make Large Language Models Generate 100% Valid Molecules?

2025-09-27 · Wen Tao, Jing Tang, Alvin Chan, Bryan Hooi 외 arxiv

Molecule generation is key to drug discovery and materials science, enabling the design of novel compounds with specific properties. Large language models (LLMs) can learn to perform a wide range of tasks from just a few…

Drug Discovery

Can Large Language Models Understand Molecules?

2024-01-05 · Shaghayegh Sadeghi, Alan Bui, Ali Forooghi, Jianguo Lu 외

Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of chemin…

Drug DiscoveryLanguage ModellingLarge Language ModelMolecular Property Prediction+3