paper-with-me

Papers

CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments

2026-04-07 · Gustav Keppler, Moritz Gstür, Veit Hagenmeyer arxiv

The advancement of Large Language Models (LLMs) has raised concerns regarding their dual-use potential in cybersecurity. Existing evaluation frameworks overwhelmingly focus on Information Technology (IT) environments, failing to capture the constraints, and specialized protocols of Operational Technology (OT). To address this gap, we introduce CritBench, a novel framework designed to evaluate the cybersecurity capabilities of LLM agents within IEC 61850 Digital Substation environments. We assess five state-of-the-art models, including OpenAI's GPT-5 suite and open-weight models, across a corpus of 81 domain-specific tasks spanning static configuration analysis, network traffic reconnaissance, and live virtual machine interaction. To facilitate industrial protocol interaction, we develop a domain-specific tool scaffold. Our empirical results show that agents reliably execute static structured-file analysis and single-tool network enumeration, but their performance degrades on dynamic tasks. Despite demonstrating explicit, internalized knowledge of the IEC 61850 standards terminology, current models struggle with the persistent sequential reasoning and state tracking required to manipulate live systems without specialized tools. Equipping agents with our domain-specific tool scaffold significantly mitigates this operational bottleneck. Code and evaluation scripts are available at: https://github.com/GKeppler/CritBench

📄 PDF Abstract BibTeX arXiv:2604.06019

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems

2026-03-30 · Yicheng Cai, Mitchell John DeStefano, Guodong Dong, Pulkit Handa 외 arxiv

As Large Language Models (LLMs) and multi-agent AI systems are demonstrating increasing potential in cybersecurity operations, organizations, policymakers, model providers, and researchers in the AI and cybersecurity com…

LLM Cyber Evaluations Don't Capture Real-World Risk

2025-01-31 · Kamilė Lukošiūtė, Adam Swanda

Large language models (LLMs) are demonstrating increasing prowess in cybersecurity applications, creating creating inherent risks alongside their potential for strengthening defenses. In this position paper, we argue tha…

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

2024-08-15 · Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji 외

Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policymakers, model providers, and researchers i…

Ollabench: Evaluating LLMs' Reasoning for Human-centric Interdependent Cybersecurity

2024-06-11 · Tam N. Nguyen

Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management. However, e…

DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments

2025-05-31 · Chiyu Zhang, Marc-Alexandre Cote, Michael Albada, Anush Sankaran 외

Large language model (LLM) agents have shown impressive capabilities in human language comprehension and reasoning, yet their potential in cybersecurity remains underexplored. We introduce DefenderBench, a practical, ope…

Large Language Model