Black-box Attacks on Image Activity Prediction and its Natural Language Explanations
Explainable AI (XAI) methods aim to describe the decision process of deep neural networks. Early XAI methods produced visual explanations, whereas more recent techniques generate multimodal explanations that include textual information and visual representations. Visual XAI methods have been shown to be vulnerable to white-box and gray-box adversarial attacks, with an attacker having full or partial knowledge of and access to the target system. As the vulnerabilities of multimodal XAI models have not been examined, in this paper we assess for the first time the robustness to black-box attacks of the natural language explanations generated by a self-rationalizing image-based activity recognition model. We generate unrestricted, spatially variant perturbations that disrupt the association between the predictions and the corresponding explanations to mislead the model into generating unfaithful explanations. We show that we can create adversarial images that manipulate the explanations of an activity recognition model by having access only to its final output.
Code (0)
등록된 구현이 없습니다.
Tasks
Activity PredictionActivity RecognitionSimilar Papers 제목 키워드 기반
SecureSense: Defending Adversarial Attack for Secure Device-Free Human Activity Recognition
Deep neural networks have empowered accurate device-free human activity recognition, which has wide applications. Deep models can extract robust features from various sensors and generalize well even in challenging situa…
Activity RecognitionAdversarial AttackHuman Activity RecognitionPerson IdentificationDefending Black-box Skeleton-based Human Activity Classifiers
Skeletal motions have been heavily replied upon for human activity recognition (HAR). Recently, a universal vulnerability of skeleton-based HAR has been identified across a variety of classifiers and data, calling for mi…
Activity RecognitionHuman Activity RecognitionTime Series AnalysisBASAR:Black-box Attack on Skeletal Action Recognition
Skeletal motion plays a vital role in human activity recognition as either an independent data source or a complement. The robustness of skeleton-based activity recognizers has been questioned recently, which shows that …
Action RecognitionActivity RecognitionAdversarial AttackHuman Activity RecognitionUnderstanding the Vulnerability of Skeleton-based Human Activity Recognition via Black-box Attack
Human Activity Recognition (HAR) has been employed in a wide range of applications, e.g. self-driving cars, where safety and lives are at stake. Recently, the robustness of skeleton-based HAR methods have been questioned…
Activity RecognitionAdversarial AttackHuman Activity RecognitionSelf-Driving Cars+1Patch of Invisibility: Naturalistic Physical Black-Box Adversarial Attacks on Object Detectors
Adversarial attacks on deep-learning models have been receiving increased attention in recent years. Work in this area has mostly focused on gradient-based techniques, so-called "white-box" attacks, wherein the attacker …
Generative Adversarial Networkobject-detectionObject Detection