paper-with-me

Papers

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

2026-05-26 · Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu arxiv

In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.

📄 PDF Abstract BibTeX arXiv:2605.27110

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Clickbait Detection in Tweets Using Self-attentive Network

2017-10-15 · Yiwei Zhou

Clickbait detection in tweets remains an elusive challenge. In this paper, we describe the solution for the Zingel Clickbait Detector at the Clickbait Challenge 2017, which is capable of evaluating each tweet's level of …

Clickbait DetectionFeature EngineeringGeneral Classification

Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning

2026-01-17 · Chaowei Zhang, Xiansheng Luo, Zewei Zhang, Yi Zhu 외 arxiv

The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for a…

Contrastive Learning

Experience-Guided Self-Adaptive Cascaded Agents for Breast Cancer Screening and Diagnosis with Reduced Biopsy Referrals

2026-02-27 · Pramit Saha, Mohammad Alsharid, Joshua Strong, J. Alison Noble arxiv

We propose an experience-guided cascaded multi-agent framework for Breast Ultrasound Screening and Diagnosis, called BUSD-Agent, that aims to reduce diagnostic escalation and unnecessary biopsy referrals. Our framework m…

Send to which account? Evaluation of an LLM-based Scambaiting System

2025-09-10 · Hossein Siadati, Haadi Jafarian, Sima Jafarikhah arxiv

Scammers are increasingly harnessing generative AI(GenAI) technologies to produce convincing phishing content at scale, amplifying financial fraud and undermining public trust. While conventional defenses, such as detect…

LLM-guided headline rewriting for clickability enhancement without clickbait

2026-03-23 · Yehudit Aperstein, Linoy Halifa, Sagiv Bar, Alexander Apartsin arxiv

Enhancing reader engagement while preserving informational fidelity is a central challenge in controllable text generation for news media. Optimizing news headlines for reader engagement is often conflated with clickbait…

Text Generation