paper-with-me

홈 › Papers

Circumventing interpretability: How to defeat mind-readers

2022-12-21 · Lee Sharkey

The increasing capabilities of artificial intelligence (AI) systems make it ever more important that we interpret their internals to ensure that their intentions are aligned with human values. Yet there is reason to believe that misaligned artificial intelligence will have a convergent instrumental incentive to make its thoughts difficult for us to interpret. In this article, I discuss many ways that a capable AI might circumvent scalable interpretability methods and suggest a framework for thinking about these potential future risks.

📄 PDF Abstract BibTeX arXiv:2212.11415

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Q-MIND: Defeating Stealthy DoS Attacks in SDN with a Machine-learning based Defense Framework

2019-07-27 · Trung V. Phan, T M Rayhan Gias, Syed Tasnimul Islam, Truong Thu Huong 외

Software Defined Networking (SDN) enables flexible and scalable network control and management. However, it also introduces new vulnerabilities that can be exploited by attackers. In particular, low-rate and slow or stea…

Anomaly DetectionBIG-bench Machine LearningManagementQ-Learning+1

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules

2026-04-03 · Cameron Pattison, Lorenzo Manuali, Seth Lazar arxiv

Safety-trained language models routinely refuse requests for help circumventing rules. But not all rules deserve compliance. When users ask for help evading rules imposed by an illegitimate authority, rules that are deep…

Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

2018-02-01 · ICML 2018 7 · Anish Athalye, Nicholas Carlini, David Wagner

We identify obfuscated gradients, a kind of gradient masking, as a phenomenon that leads to a false sense of security in defenses against adversarial examples. While defenses that cause obfuscated gradients appear to def…

Adversarial AttackAdversarial Defense

TStarBots: Defeating the Cheating Level Builtin AI in StarCraft II in the Full Game

2018-09-19 · Peng Sun, Xinghai Sun, Lei Han, Jiechao Xiong 외

Starcraft II (SC2) is widely considered as the most challenging Real Time Strategy (RTS) game. The underlying challenges include a large observation space, a huge (continuous and infinite) action space, partial observati…

AI AgentDecision MakingDeep Reinforcement LearningReal-Time Strategy Games+3

Defeat Devices in AI Systems

2026-06-27 · Emilio Ferrara arxiv

AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagging, benchmark gaming, deceptive scheming, specification gaming, and trojans have…