paper-with-me

홈 › Papers

Well, that escalated quickly: The Single-Turn Crescendo Attack (STCA)

2024-09-04 · Alan Aqrawi, Arian Abbasi

This paper introduces a new method for adversarial attacks on large language models (LLMs) called the Single-Turn Crescendo Attack (STCA). Building on the multi-turn crescendo attack method introduced by Russinovich, Salem, and Eldan (2024), which gradually escalates the context to provoke harmful responses, the STCA achieves similar outcomes in a single interaction. By condensing the escalation into a single, well-crafted prompt, the STCA bypasses typical moderation filters that LLMs use to prevent inappropriate outputs. This technique reveals vulnerabilities in current LLMs and emphasizes the importance of stronger safeguards in responsible AI (RAI). The STCA offers a novel method that has not been previously explored.

📄 PDF Abstract BibTeX arXiv:2409.03131

Code (1)

alanaqrawi/stca 공식 구현

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

2024-04-02 · Mark Russinovich, Ahmed Salem, Ronen Eldan

Large Language Models (LLMs) have risen significantly in popularity and are increasingly being adopted across multiple applications. These LLMs are heavily aligned to resist engaging in illegal or unethical topics as a m…

LLM Jailbreak

An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)

2024-11-27 · Ted Kwartler, Nataliia Bagan, Ivan Banny, Alan Aqrawi 외

The Single-Turn Crescendo Attack (STCA), first introduced in Aqrawi and Abbasi [2024], is an innovative method designed to bypass the ethical safeguards of text-to-text AI models, compelling them to generate harmful cont…

CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior

2017-10-30 · ICLR 2018 1 · Xiang Zhang, Nishant Vishwamitra, Hongxin Hu, Feng Luo

We introduce a new deep convolutional neural network, CrescendoNet, by stacking simple building blocks without residual connections. Each Crescendo block contains independent convolution paths with increased depths. The …

Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search

2025-03-13 · Andy Zhou, Ron Arel

We introduce Tempest, a multi-turn adversarial framework that models the gradual erosion of Large Language Model (LLM) safety through a tree search perspective. Unlike single-turn jailbreaks that rely on one meticulously…

Language ModelingLanguage ModellingLarge Language Model

AJAR: Adaptive Jailbreak Architecture for Red-teaming

2026-01-16 · Yipu Dou, Wang Yang arxiv

Large language model (LLM) safety evaluation is moving from content moderation to action security as modern systems gain persistent state, tool access, and autonomous control loops. Existing jailbreak frameworks still le…