Don't Lie to Me: Avoiding Malicious Explanations with STEALTH
STEALTH is a method for using some AI-generated model, without suffering from malicious attacks (i.e. lying) or associated unfairness issues. After recursively bi-clustering the data, STEALTH system asks the AI model a limited number of queries about class labels. STEALTH asks so few queries (1 per data cluster) that malicious algorithms (a) cannot detect its operation, nor (b) know when to lie.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringSimilar Papers 제목 키워드 기반
Analyzing Federated Learning through an Adversarial Lens
Federated learning distributes model training among a multitude of agents, who, guided by privacy concerns, perform training using their local data but share only model parameter updates, for iterative aggregation at the…
Federated LearningModel Poisoningparameter estimationFool SHAP with Stealthily Biased Sampling
SHAP explanations aim at identifying which features contribute the most to the difference in model prediction at a specific input versus a background distribution. Recent studies have shown that they can be manipulated b…
FairnessStealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
The integration of large language models (LLMs) into information retrieval systems introduces new attack surfaces, particularly for adversarial ranking manipulations. We present StealthRank, a novel adversarial ranking a…
Adversarial TextInformation RetrievalProduct RecommendationRecommendation SystemsCan Quantum Federated Learning Withstand Circuit-Level Backdoors?
Quantum Federated Learning (QFL) inherits the core vulnerability of federated optimization to malicious clients, while also introducing an attack surface from variational circuit training and measurement-driven gradients…
Federated LearningStealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
Large language model (LLM) safety is a critical issue, with numerous studies employing red team testing to enhance model security. Among these, jailbreak methods explore potential vulnerabilities by crafting malicious pr…
Language ModelingLanguage ModellingLarge Language Model