paper-with-me

홈 › Papers

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

2026-02-12 · Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin arxiv

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments. The benchmark and code are available at https://github.com/causalNLP/gt-harmbench.

📄 PDF Abstract BibTeX arXiv:2602.12316

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

2024-02-06 · Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 외

Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigoro…

Red Teaming

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

2026-04-27 · Xinran Zhang arxiv

Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combination of judge model and judge prompt, is typically treated as a fixed …

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

2026-07-01 · Tong Xu, Xinzhe Cao, Zhihui Zhu, Keyan Ding 외 arxiv

Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, ma…

RedDebate: Safer Responses through Multi-Agent Red Teaming Debates

2025-06-04 · Ali Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan Zhu

We propose RedDebate, a novel multi-agent debate framework that leverages adversarial argumentation among Large Language Models (LLMs) to proactively identify and mitigate their own unsafe behaviours. Existing AI safety …

Red Teaming

A Safety and Security Framework for Real-World Agentic Systems

2025-11-27 · Shaona Ghosh, Barnaby Simkin, Kyriacos Shiarlis, Soumili Nandi 외 arxiv

This paper introduces a dynamic and actionable framework for securing agentic AI systems in enterprise deployment. We contend that safety and security are not merely fixed attributes of individual models but also emergen…

Red Teaming