paper-with-me

홈 › Papers

From Feature Visualization to Visual Circuits: Effect of Adversarial Model Manipulation

2024-06-03 · Geraldin Nanfack, Michael Eickenberg, Eugene Belilovsky

Understanding the inner working functionality of large-scale deep neural networks is challenging yet crucial in several high-stakes applications. Mechanistic inter- pretability is an emergent field that tackles this challenge, often by identifying human-understandable subgraphs in deep neural networks known as circuits. In vision-pretrained models, these subgraphs are usually interpreted by visualizing their node features through a popular technique called feature visualization. Recent works have analyzed the stability of different feature visualization types under the adversarial model manipulation framework. This paper starts by addressing limitations in existing works by proposing a novel attack called ProxPulse that simultaneously manipulates the two types of feature visualizations. Surprisingly, when analyzing these attacks under the umbrella of visual circuits, we find that visual circuits show some robustness to ProxPulse. We, therefore, introduce a new attack based on ProxPulse that unveils the manipulability of visual circuits, shedding light on their lack of robustness. The effectiveness of these attacks is validated using pre-trained AlexNet and ResNet-50 models on ImageNet.

📄 PDF Abstract BibTeX arXiv:2406.01365

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Don't trust your eyes: on the (un)reliability of feature visualizations

2023-06-07 · Robert Geirhos, Roland S. Zimmermann, Blair Bilodeau, Wieland Brendel 외

How do neural networks extract patterns from pixels? Feature visualizations attempt to answer this important question by visualizing highly activating patterns through optimization. Today, visualization methods form the …

Analyzing the Noise Robustness of Deep Neural Networks

2020-01-26 · Kelei Cao, Mengchen Liu, Hang Su, Jing Wu 외

Adversarial examples, generated by adding small but intentionally imperceptible perturbations to normal examples, can mislead deep neural networks (DNNs) to make incorrect predictions. Although much work has been done on…

Adversarial Attack

VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow

2025-03-28 · Ada Gorgun, Bernt Schiele, Jonas Fischer

Neural networks are widely adopted to solve complex and challenging tasks. Especially in high-stakes decision-making, understanding their reasoning process is crucial, yet proves challenging for modern deep networks. Fea…

Decision Making

Efficient Visualization of Neural Networks with Generative Models and Adversarial Perturbations

2024-09-20 · Athanasios Karagounis

This paper presents a novel approach for deep visualization via a generative network, offering an improvement over existing methods. Our model simplifies the architecture by reducing the number of networks used, requirin…

Image Generation

Adversarial Attacks on Machine Learning-Aided Visualizations

2024-09-04 · Takanori Fujiwara, Kostiantyn Kucher, Junpeng Wang, Rafael M. Martins 외

Research in ML4VIS investigates how to use machine learning (ML) techniques to generate visualizations, and the field is rapidly growing with high societal impact. However, as with any computational pipeline that employs…