paper-with-me

홈 › Papers

The Distillation Game: Adaptive Attacks & Efficient Defenses

2026-05-21 · Youssef Allouah, Mahdi Haghifam, Sanmi Koyejo, Reza Shokri arxiv

Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-constrained teacher and an adaptive student. Our framework yields tractable one-sided response rules: an adaptive evaluation rule in which the student reweights high-value examples, and a teacher-side defense template that suppresses outputs most useful for distillation. From a cheap proxy for example value, we derive Product-of-Experts (PoE), a simple forward-pass-only defense that combines the teacher with a proxy student during generation. Empirically, adaptive evaluation reveals a large passive--adaptive gap: on state-of-the-art defenses, adaptive students recover substantially more capability than passive evaluation suggests on GSM8K and MATH. Under this stronger evaluation, the apparent robustness gap between expensive defenses and PoE narrows considerably, while PoE remains substantially cheaper and preserves higher-quality reasoning traces. Overall, our results suggest that strong distillation remains difficult to stop, and that progress on antidistillation should be judged against adaptive students rather than passive ones. Our code is available at: https://github.com/ysfalh/distillation-game.

📄 PDF Abstract BibTeX arXiv:2605.22737

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Adaptive Arms Race: Redefining Robustness in AI Security

2023-12-20 · Ilias Tsingenopoulos, Vera Rimmer, Davy Preuveneers, Fabio Pierazzi 외

Despite considerable efforts on making them robust, real-world AI-based systems remain vulnerable to decision based attacks, as definitive proofs of their operational robustness have so far proven intractable. Canonical …

Beware the Black-Box: on the Robustness of Recent Defenses to Adversarial Examples

2020-06-18 · Kaleel Mahmood, Deniz Gurevin, Marten van Dijk, Phuong Ha Nguyen

Many defenses have recently been proposed at venues like NIPS, ICML, ICLR and CVPR. These defenses are mainly focused on mitigating white-box attacks. They do not properly examine black-box attacks. In this paper, we exp…

Diversity

A Game Theoretic Analysis of Additive Adversarial Attacks and Defenses

2020-09-14 · NeurIPS 2020 12 · Ambar Pal, René Vidal

Research in adversarial learning follows a cat and mouse game between attackers and defenders where attacks are proposed, they are mitigated by new defenses, and subsequently new attacks are proposed that break earlier d…

On Certifying Robustness against Backdoor Attacks via Randomized Smoothing

2020-02-26 · Binghui Wang, Xiaoyu Cao, Jinyuan Jia, Neil Zhenqiang Gong

Backdoor attack is a severe security threat to deep neural networks (DNNs). We envision that, like adversarial examples, there will be a cat-and-mouse game for backdoor attacks, i.e., new empirical defenses are developed…

Backdoor Attack

Game Theoretic Mixed Experts for Combinational Adversarial Machine Learning

2022-11-26 · Ethan Rathbun, Kaleel Mahmood, Sohaib Ahmad, Caiwen Ding 외

Recent advances in adversarial machine learning have shown that defenses considered to be robust are actually susceptible to adversarial attacks which are specifically customized to target their weaknesses. These defense…

Adversarial Defense