paper-with-me

Papers

A Framework for Evaluating Emerging Cyberattack Capabilities of AI

2025-03-14 · Mikel Rodriguez, Raluca Ada Popa, Four Flynn, Lihao Liang, Allan Dafoe, Anna Wang

As frontier AI models become more capable, evaluating their potential to enable cyberattacks is crucial for ensuring the safe development of Artificial General Intelligence (AGI). Current cyber evaluation efforts are often ad-hoc, lacking systematic analysis of attack phases and guidance on targeted defenses. This work introduces a novel evaluation framework that addresses these limitations by: (1) examining the end-to-end attack chain, (2) identifying gaps in AI threat evaluation, and (3) helping defenders prioritize targeted mitigations and conduct AI-enabled adversary emulation for red teaming. Our approach adapts existing cyberattack chain frameworks for AI systems. We analyzed over 12,000 real-world instances of AI involvement in cyber incidents, catalogued by Google's Threat Intelligence Group, to curate seven representative attack chain archetypes. Through a bottleneck analysis on these archetypes, we pinpointed phases most susceptible to AI-driven disruption. We then identified and utilized externally developed cybersecurity model evaluations focused on these critical phases. We report on AI's potential to amplify offensive capabilities across specific attack stages, and offer recommendations for prioritizing defenses. We believe this represents the most comprehensive AI cyber risk evaluation framework published to date.

📄 PDF Abstract BibTeX arXiv:2503.11917

Code (0)

등록된 구현이 없습니다.

Tasks

Red Teaming

Similar Papers 제목 키워드 기반

Evaluating the Cybersecurity Risk of Real World, Machine Learning Production Systems

2021-07-05 · Ron Bitton, Nadav Maman, Inderjeet Singh, Satoru Momiyama 외

Although cyberattacks on machine learning (ML) production systems can be harmful, today, security practitioners are ill equipped, lacking methodologies and tactical tools that would allow them to analyze the security ris…

BIG-bench Machine LearningGraph Generation

When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs

2024-10-18 · Hanna Kim, Minkyoo Song, Seung Ho Na, Seungwon Shin 외

Recent advancements in Large Language Models (LLMs) have established them as agentic systems capable of planning and interacting with various tools. These LLM agents are often paired with web-based tools, enabling access…

Defense against DoS and load altering attacks via model-free control: A proposal for a new cybersecurity setting

2021-07-09 · Michel Fliess, Cédric Join, Dominique Sauter

Defense against cyberattacks is an emerging topic related to fault-tolerant control. In order to avoid difficult mathematical modeling, model-free control (MFC) is suggested as an alternative to classical control. For il…

A Dual-Tier Adaptive One-Class Classification IDS for Emerging Cyberthreats

2024-03-17 · Md. Ashraf Uddin, Sunil Aryal, Mohamed Reda Bouadjenek, Muna Al-Hawawreh 외

In today's digital age, our dependence on IoT (Internet of Things) and IIoT (Industrial IoT) systems has grown immensely, which facilitates sensitive activities such as banking transactions and personal, enterprise data,…

ClusteringIntrusion DetectionNetwork Intrusion DetectionOne-Class Classification

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

2026-04-21 · Euntae Kim, Soomin Han, Buru Chang arxiv

Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to complete, revise, and refine their content. However, this capability pose…