paper-with-me

홈 › Papers

Feature-Guided Black-Box Safety Testing of Deep Neural Networks

2017-10-21 · Matthew Wicker, Xiaowei Huang, Marta Kwiatkowska

Despite the improved accuracy of deep neural networks, the discovery of adversarial examples has raised serious safety concerns. Most existing approaches for crafting adversarial examples necessitate some knowledge (architecture, parameters, etc.) of the network at hand. In this paper, we focus on image classifiers and propose a feature-guided black-box approach to test the safety of deep neural networks that requires no such knowledge. Our algorithm employs object detection techniques such as SIFT (Scale Invariant Feature Transform) to extract features from an image. These features are converted into a mutable saliency distribution, where high probability is assigned to pixels that affect the composition of the image with respect to the human visual system. We formulate the crafting of adversarial examples as a two-player turn-based stochastic game, where the first player's objective is to minimise the distance to an adversarial example by manipulating the features, and the second player can be cooperative, adversarial, or random. We show that, theoretically, the two-player game can con- verge to the optimal strategy, and that the optimal strategy represents a globally minimal adversarial image. For Lipschitz networks, we also identify conditions that provide safety guarantees that no adversarial examples exist. Using Monte Carlo tree search we gradually explore the game state space to search for adversarial examples. Our experiments show that, despite the black-box setting, manipulations guided by a perception-based saliency distribution are competitive with state-of-the-art methods that rely on white-box saliency matrices or sophisticated optimization procedures. Finally, we show how our method can be used to evaluate robustness of neural networks in safety-critical applications such as traffic sign recognition in self-driving cars.

📄 PDF Abstract BibTeX arXiv:1710.07859

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject DetectionSelf-Driving CarsTraffic Sign Recognition

Similar Papers 제목 키워드 기반

Testing Neural Networks via Bayesian-Guided Exploration of Decision Landscapes

2026-06-03 · Bin Duan, Meiru Che, Guowei Yang arxiv

As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability. Existing testing methods, whether black-box or white-box, primarily use global muta…

PAL: Proxy-Guided Black-Box Attack on Large Language Models

2024-02-15 · Chawin Sitawarin, Norman Mu, David Wagner, Alexandre Araujo

Large Language Models (LLMs) have surged in popularity in recent months, but they have demonstrated concerning capabilities to generate harmful content when manipulated. While techniques like safety fine-tuning aim to mi…

Efficient Safety Testing of Autonomous Vehicles via Adaptive Search over Crash-Derived Scenarios

2025-08-07 · Rui Zhou arxiv

Ensuring the safety of autonomous vehicles (AVs) is paramount in their development and deployment. Safety-critical scenarios pose more severe challenges, necessitating efficient testing methods to validate AVs safety. Th…

Autonomous Vehicles

Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation

2025-12-05 · Mara Downing, Matthew Peng, Jacob Granley, Michael Beyeler 외 arxiv

Objective: Machine learning (ML) models are increasingly used to generate electrical stimulation patterns in neuroprosthetic devices such as visual prostheses. While these models promise precise and personalized control,…

Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models

2024-08-01 · Yingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li 외

Text-to-image (T2I) generative models have revolutionized content creation by transforming textual descriptions into high-quality images. However, these models are vulnerable to jailbreaking attacks, where carefully craf…

Image GenerationIn-Context LearningLanguage ModellingLarge Language Model+2