paper-with-me

홈 › Papers

Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks

2025-10-16 · ChenYu Wu, Yi Wang, Yang Liao arxiv

Large language models (LLMs) are increasingly vulnerable to multi-turn jailbreak attacks, where adversaries iteratively elicit harmful behaviors that bypass single-turn safety filters. Existing defenses predominantly rely on passive rejection, which either fails against adaptive attackers or overly restricts benign users. We propose a honeypot-based proactive guardrail system that transforms risk avoidance into risk utilization. Our framework fine-tunes a bait model to generate ambiguous, non-actionable but semantically relevant responses, which serve as lures to probe user intent. Combined with the protected LLM's safe reply, the system inserts proactive bait questions that gradually expose malicious intent through multi-turn interactions. We further introduce the Honeypot Utility Score (HUS), measuring both the attractiveness and feasibility of bait responses, and use a Defense Efficacy Rate (DER) for balancing safety and usability. Initial experiment on MHJ Datasets with recent attack method across GPT-4o show that our system significantly disrupts jailbreak success while preserving benign user experience.

📄 PDF Abstract BibTeX arXiv:2510.15017

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM Honeypot: Leveraging Large Language Models as Advanced Interactive Honeypot Systems

2024-09-12 · Hakan T. Otal, M. Abdullah Canbaz

The rapid evolution of cyber threats necessitates innovative solutions for detecting and analyzing malicious activity. Honeypots, which are decoy systems designed to lure and interact with attackers, have emerged as a cr…

Language ModelingLanguage ModellingModel SelectionPrompt Engineering

What are Attackers after on IoT Devices? An approach based on a multi-phased multi-faceted IoT honeypot ecosystem and data clustering

2021-12-21 · Armin Ziaie Tabari, Xinming Ou, Anoop Singhal

The growing number of Internet of Things (IoT) devices makes it imperative to be aware of the real-world threats they face in terms of cybersecurity. While honeypots have been historically used as decoy devices to help r…

Measuring and Clustering Network Attackers using Medium-Interaction Honeypots

2022-06-27 · Zain Shamsi, Daniel Zhang, Daehyun Kyoung, Alex Liu

Network honeypots are often used by information security teams to measure the threat landscape in order to secure their networks. With the advancement of honeypot development, today's medium-interaction honeypots provide…

Clustering

Social Honeypot for Humans: Luring People through Self-managed Instagram Pages

2023-03-31 · Sara Bardi, Mauro Conti, Luca Pajola, Pier Paolo Tricomi

Social Honeypots are tools deployed in Online Social Networks (OSN) to attract malevolent activities performed by spammers and bots. To this end, their content is designed to be of maximum interest to malicious users. Ho…

Marketing

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

2026-05-28 · Mark Vero, Fabian Kaczmarczyck, Ivan Petrov, Ilia Shumailov 외 arxiv

Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-inte…