paper-with-me

Papers

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

2025-10-13 · Zicheng Liu, Lige Huang, Jie Zhang, Dongrui Liu, Yuan Tian, Jing Shao arxiv

The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PACEbench, a practical AI cyber-exploitation benchmark built on the principles of realistic vulnerability difficulty, environmental complexity, and cyber defenses. Specifically, PACEbench comprises four scenarios spanning single, blended, chained, and defense vulnerability exploitations. To handle these complex challenges, we propose PACEagent, a novel agent that emulates human penetration testers by supporting multi-phase reconnaissance, analysis, and exploitation. Extensive experiments with seven frontier LLMs demonstrate that current models struggle with complex cyber scenarios, and none can bypass defenses. These findings suggest that current models do not yet pose a generalized cyber offense threat. Nonetheless, our work provides a robust benchmark to guide the trustworthy development of future models.

📄 PDF Abstract BibTeX arXiv:2510.11688

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI Cyber Risk Benchmark: Automated Exploitation Capabilities

2024-10-29 · Dan Ristea, Vasilios Mavroudis, Chris Hicks

We introduce a new benchmark for assessing AI models' capabilities and risks in automated software exploitation, focusing on their ability to detect and exploit vulnerabilities in real-world software systems. Using DARPA…

BenchmarkingVulnerability Detection

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges

2026-06-12 · Fengyu Liu, Jiarun Dai, Yihe Fan, Wuyuao Mai 외 arxiv

Frontier AI systems are increasingly capable of cybersecurity tasks, including codebase inspection, vulnerability detection, and exploitation. However, evaluating their offensive capabilities remains constrained by limit…

Vulnerability Detection

Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

2024-10-10 · Andrey Anurin, Jonathan Ng, Kibo Schaffer, Jason Schreiber 외

LLM agents have the potential to revolutionize defensive cyber operations, but their offensive capabilities are not yet fully understood. To prepare for emerging threats, model developers and governments are evaluating t…

HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities

2025-10-14 · Xiaoxue Ren, Penghao Jiang, Kaixin Li, Zhiyong Huang 외 arxiv

Web applications are prime targets for cyberattacks as gateways to critical services and sensitive data. Traditional penetration testing is costly and expertise-intensive, making it difficult to scale with the growing we…

Vulnerability Detection

Enabling Cyber Security Education through Digital Twins and Generative AI

2025-07-23 · Vita Santa Barletta, Vito Bavaro, Miriana Calvano, Antonio Curci 외 arxiv

Digital Twins (DTs) are gaining prominence in cybersecurity for their ability to replicate complex IT (Information Technology), OT (Operational Technology), and IoT (Internet of Things) infrastructures, allowing for real…