paper-with-me

홈 › Papers

Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

2024-10-10 · Andrey Anurin, Jonathan Ng, Kibo Schaffer, Jason Schreiber, Esben Kran

LLM agents have the potential to revolutionize defensive cyber operations, but their offensive capabilities are not yet fully understood. To prepare for emerging threats, model developers and governments are evaluating the cyber capabilities of foundation models. However, these assessments often lack transparency and a comprehensive focus on offensive capabilities. In response, we introduce the Catastrophic Cyber Capabilities Benchmark (3CB), a novel framework designed to rigorously assess the real-world offensive capabilities of LLM agents. Our evaluation of modern LLMs on 3CB reveals that frontier models, such as GPT-4o and Claude 3.5 Sonnet, can perform offensive tasks such as reconnaissance and exploitation across domains ranging from binary analysis to web technologies. Conversely, smaller open-source models exhibit limited offensive capabilities. Our software solution and the corresponding benchmark provides a critical tool to reduce the gap between rapidly improving capabilities and robustness of cyber offense evaluations, aiding in the safer deployment and regulation of these powerful technologies.

📄 PDF Abstract BibTeX arXiv:2410.09114

Code (1)

apartresearch/3cb 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems

2026-03-30 · Yicheng Cai, Mitchell John DeStefano, Guodong Dong, Pulkit Handa 외 arxiv

As Large Language Models (LLMs) and multi-agent AI systems are demonstrating increasing potential in cybersecurity operations, organizations, policymakers, model providers, and researchers in the AI and cybersecurity com…

CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence

2024-06-11 · Md Tanvirul Alam, Dipkamal Bhusal, Le Nguyen, Nidhi Rastogi

Cyber threat intelligence (CTI) is crucial in today's cybersecurity landscape, providing essential insights to understand and mitigate the ever-evolving cyber threats. The recent rise of Large Language Models (LLMs) have…

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

2025-10-13 · Zicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu 외 arxiv

The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accuratel…

AI Cyber Risk Benchmark: Automated Exploitation Capabilities

2024-10-29 · Dan Ristea, Vasilios Mavroudis, Chris Hicks

We introduce a new benchmark for assessing AI models' capabilities and risks in automated software exploitation, focusing on their ability to detect and exploit vulnerabilities in real-world software systems. Using DARPA…

BenchmarkingVulnerability Detection

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges

2026-06-12 · Fengyu Liu, Jiarun Dai, Yihe Fan, Wuyuao Mai 외 arxiv

Frontier AI systems are increasingly capable of cybersecurity tasks, including codebase inspection, vulnerability detection, and exploitation. However, evaluating their offensive capabilities remains constrained by limit…

Vulnerability Detection