paper-with-me

Papers

Learning the Wrong Lessons: Inserting Trojans During Knowledge Distillation

2023-03-09 · Leonard Tang, Tom Shlomi, Alexander Cai

In recent years, knowledge distillation has become a cornerstone of efficiently deployed machine learning, with labs and industries using knowledge distillation to train models that are inexpensive and resource-optimized. Trojan attacks have contemporaneously gained significant prominence, revealing fundamental vulnerabilities in deep learning models. Given the widespread use of knowledge distillation, in this work we seek to exploit the unlabelled data knowledge distillation process to embed Trojans in a student model without introducing conspicuous behavior in the teacher. We ultimately devise a Trojan attack that effectively reduces student accuracy, does not alter teacher performance, and is efficiently constructible in practice.

📄 PDF Abstract BibTeX arXiv:2303.05593

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

Hu-Fu: Hardware and Software Collaborative Attack Framework against Neural Networks

2018-05-14 · Wenshuo Li, Jincheng Yu, Xuefei Ning, Pengjun Wang 외

Recently, Deep Learning (DL), especially Convolutional Neural Network (CNN), develops rapidly and is applied to many tasks, such as image classification, face recognition, image segmentation, and human detection. Due to …

Autonomous DrivingCloud ComputingFace RecognitionGeneral Classification+5

PoTrojan: powerful neural-level trojan designs in deep learning models

2018-02-08 · Minhui Zou, Yang Shi, Chengliang Wang, Fangyu Li 외

With the popularity of deep learning (DL), artificial intelligence (AI) has been applied in many areas of human life. Neural network or artificial neural network (NN), the main technique behind DL, has been extensively s…

Deep Learning

Trojans in Artificial Intelligence (TrojAI) Final Report

2026-02-06 · Kristopher W. Reese, Taylor Kulp-McDowall, Michael Majurski, Tim Blattner 외 arxiv

The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI Trojans. These AI trojans are malicious, …

Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing

2024-12-17 · Keltin Grimes, Marco Christiani, David Shriver, Marissa Connor

Model editing methods modify specific behaviors of Large Language Models by altering a small, targeted set of network weights and require very little data and compute. These methods can be used for malicious applications…

MisinformationModel Editing

Trojan Attacks on Wireless Signal Classification with Adversarial Machine Learning

2019-10-23 · Kemal Davaslioglu, Yalin E. Sagduyu

We present a Trojan (backdoor or trapdoor) attack that targets deep learning applications in wireless communications. A deep learning classifier is considered to classify wireless signals using raw (I/Q) samples as featu…

BIG-bench Machine LearningClassificationClusteringDeep Learning+2