paper-with-me

Papers

A Sweet Rabbit Hole by DARCY: Using Honeypots to Detect Universal Trigger's Adversarial Attacks

2020-11-20 · ACL 2021 5 · Thai Le, Noseong Park, Dongwon Lee

The Universal Trigger (UniTrigger) is a recently-proposed powerful adversarial textual attack method. Utilizing a learning-based mechanism, UniTrigger generates a fixed phrase that, when added to any benign inputs, can drop the prediction accuracy of a textual neural network (NN) model to near zero on a target class. To defend against this attack that can cause significant harm, in this paper, we borrow the "honeypot" concept from the cybersecurity community and propose DARCY, a honeypot-based defense framework against UniTrigger. DARCY greedily searches and injects multiple trapdoors into an NN model to "bait and catch" potential attacks. Through comprehensive experiments across four public datasets, we show that DARCY detects UniTrigger's adversarial attacks with up to 99% TPR and less than 2% FPR in most cases, while maintaining the prediction accuracy (in F1) for clean inputs within a 1% margin. We also demonstrate that DARCY with multiple trapdoors is also robust to a diverse set of attack scenarios with attackers' varying levels of knowledge and skills. Source code will be released upon the acceptance of this paper.

📄 PDF Abstract BibTeX arXiv:2011.10492

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial Attack

Similar Papers 제목 키워드 기반

Black Holes and White Rabbits: Metaphor Identification with Visual Features

2016-06-01 · NAACL 2016 6 · Ekaterina Shutova, Douwe Kiela, Jean Maillard
Clustering

Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots

2026-05-28 · Mark Vero, Fabian Kaczmarczyck, Ivan Petrov, Ilia Shumailov 외 arxiv

Honeypots are decoy systems mimicking real system components designed to defend against cyber attacks. Recently, LLMs increasingly serve as simulation backbones for honeypots. They enable defenders to construct high-inte…

Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models

2023-09-08 · Arka Dutta, Adel Khorramrouz, Sujan Dutta, Ashiqur R. KhudaBukhsh

This paper makes three contributions. First, it presents a generalizable, novel framework dubbed \textit{toxicity rabbit hole} that iteratively elicits toxic content from a wide suite of large language models. Spanning a…

Machine Learning for Detection and Severity Estimation of Sweetpotato Weevil Damage in Field and Lab Conditions

2026-02-06 · Doreen M. Chelangat, Sudi Murindanyi, Bruce Mugizi, Paul Musana 외 arxiv

Sweetpotato weevils (Cylas spp.) are considered among the most destructive pests impacting sweetpotato production, particularly in sub-Saharan Africa. Traditional methods for assessing weevil damage, predominantly relyin…

Object Detection

Security Orchestration, Automation, and Response Engine for Deployment of Behavioural Honeypots

2022-01-14 · Upendra Bartwal, Subhasis Mukhopadhyay, Rohit Negi, Sandeep Shukla

Cyber Security is a critical topic for organizations with IT/OT networks as they are always susceptible to attack, whether insider or outsider. Since the cyber landscape is an ever-evolving scenario, one must keep upgrad…

Intrusion DetectionManagement