paper-with-me

Papers

Exploring the Adversarial Frontier: Quantifying Robustness via Adversarial Hypervolume

2024-03-08 · Ping Guo, Cheng Gong, Xi Lin, Zhiyuan Yang, Qingfu Zhang

The escalating threat of adversarial attacks on deep learning models, particularly in security-critical fields, has underscored the need for robust deep learning systems. Conventional robustness evaluations have relied on adversarial accuracy, which measures a model's performance under a specific perturbation intensity. However, this singular metric does not fully encapsulate the overall resilience of a model against varying degrees of perturbation. To address this gap, we propose a new metric termed adversarial hypervolume, assessing the robustness of deep learning models comprehensively over a range of perturbation intensities from a multi-objective optimization standpoint. This metric allows for an in-depth comparison of defense mechanisms and recognizes the trivial improvements in robustness afforded by less potent defensive strategies. Additionally, we adopt a novel training algorithm that enhances adversarial robustness uniformly across various perturbation intensities, in contrast to methods narrowly focused on optimizing adversarial accuracy. Our extensive empirical studies validate the effectiveness of the adversarial hypervolume metric, demonstrating its ability to reveal subtle differences in robustness that adversarial accuracy overlooks. This research contributes a new measure of robustness and establishes a standard for assessing and benchmarking the resilience of current and future defensive models against adversarial threats.

📄 PDF Abstract BibTeX arXiv:2403.05100

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessBenchmarkingDeep Learning

Similar Papers 제목 키워드 기반

Quantifying Assistive Robustness Via the Natural-Adversarial Frontier

2023-10-16 · Jerry Zhi-Yang He, Zackory Erickson, Daniel S. Brown, Anca D. Dragan

Our ultimate goal is to build robust policies for robots that assist people. What makes this hard is that people can behave unexpectedly at test time, potentially interacting with the robot outside its training distribut…

Scaling Trends in Language Model Robustness

2024-07-25 · Nikolaus Howe, Ian McKenzie, Oskar Hollinsworth, Michał Zajac 외

Increasing model size has unlocked a dazzling array of capabilities in modern language models. At the same time, even frontier models remain vulnerable to jailbreaks and prompt injections, despite concerted efforts to ma…

Adversarial RobustnessLanguage ModelingLanguage Modellingmodel

Advancing NLP Security by Leveraging LLMs as Adversarial Engines

2024-10-23 · Sudarshan Srinivasan, Maria Mahbub, Amir Sadovnik

This position paper proposes a novel approach to advancing NLP security by leveraging Large Language Models (LLMs) as engines for generating diverse adversarial attacks. Building upon recent work demonstrating LLMs' effe…

Position

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

2026-04-20 · Marcello Galisai, Susanna Cifani, Francesco Giarrusso, Piercosma Bisconti 외 arxiv

The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewri…

MixAT: Combining Continuous and Discrete Adversarial Training for LLMs

2025-05-22 · Csaba Dékány, Stefan Balauca, Robin Staab, Dimitar I. Dimitrov 외

Despite recent efforts in Large Language Models (LLMs) safety and alignment, current adversarial attacks on frontier LLMs are still able to force harmful generations consistently. Although adversarial training has been w…