paper-with-me

Papers

Improving Alignment and Robustness with Circuit Breakers

2024-06-06 · Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, Dan Hendrycks

AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interrupts the models as they respond with harmful outputs with "circuit breakers." Existing techniques aimed at improving alignment, such as refusal training, are often bypassed. Techniques such as adversarial training try to plug these holes by countering specific attacks. As an alternative to refusal training and adversarial training, circuit-breaking directly controls the representations that are responsible for harmful outputs in the first place. Our technique can be applied to both text-only and multimodal language models to prevent the generation of harmful outputs without sacrificing utility -- even in the presence of powerful unseen attacks. Notably, while adversarial robustness in standalone image recognition remains an open challenge, circuit breakers allow the larger multimodal system to reliably withstand image "hijacks" that aim to produce harmful content. Finally, we extend our approach to AI agents, demonstrating considerable reductions in the rate of harmful actions when they are under attack. Our approach represents a significant step forward in the development of reliable safeguards to harmful behavior and adversarial attacks.

📄 PDF Abstract BibTeX arXiv:2406.04313

Code (4)

blackswan-ai/circuit-breakers 공식 구현 pytorch
blackswan-ai/short-circuiting 공식 구현 pytorch
grayswanai/circuit-breakers 공식 구현 pytorch
schwinnl/circuit-breakers-eval pytorch

Tasks

Adversarial Robustness

Similar Papers 제목 키워드 기반

A Risk-Based Probabilistic Transient Stability Approach for Ranking of Circuit Breakers in a Power System

2025-05-21 · Umair Shahzad

Power systems are getting more complex than ever and are consequently operating close to their limit of stability. Moreover, with the increasing demand of renewable wind generation, and the requirement to maintain a secu…

Reducing the Scope of Language Models

2024-10-28 · David Yunis, Siyu Huo, Chulaka Gunasekara, Danish Contractor

We now deploy language models in a wide variety of user-facing applications. Typically, these deployments have some specific purpose, like answering questions about documentation or acting as coding assistants, but they …

DiversitySentiment Analysis

Multi-agent Deep Reinforcement Learning for Distributed Load Restoration

2023-06-24 · Linh Vu, Tuyen Vu, Thanh-Long Vu, Anurag Srivastava

This paper addresses the load restoration problem after power outage events. Our primary proposed methodology is using multi-agent deep reinforcement learning to optimize the load restoration process in distribution syst…

Deep Reinforcement Learningreinforcement-learningReinforcement Learning

Rulebreakers Challenge: Revealing a Blind Spot in Large Language Models' Reasoning with Formal Logic

2024-10-21 · Jason Chan, Robert Gaizauskas, Zhixue Zhao

Formal logic has long been applied to natural language reasoning, but this approach can sometimes lead to conclusions that, while logically entailed, are factually inconsistent with the premises or are not typically infe…

Formal LogicWorld Knowledge

Identifying Dealbreakers and Robust Policies for the Energy Transition Amid Unexpected Events

2025-02-19 · Diederik Coppitters, Gabriel Wiest, Leonard Göke, Francesco Contino 외

Disruptions in energy imports, backlash in social acceptance, and novel technologies failing to develop are unexpected events that are often overlooked in energy planning, despite their ability to jeopardize the energy t…