paper-with-me

홈 › Papers

CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

2026-05-08 · Taein Lim, Seongyong Ju, Munhyeok Kim, Hyunjun Kim, Hoki Kim arxiv

Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically, each agent exhibits an attack-selection bias, disproportionately concentrating its efforts on a narrow subset of attack families regardless of prompt variations. To systematically quantify this behavior, we introduce CyBiasBench, a comprehensive 630-session benchmark that evaluates five agents on three targets and four prompt conditions with ten attack families. We identify explicit bias across agents, with different dominant attack families and varying entropy levels in their attack-family allocation distributions. Such bias is better characterized as a trait of the agents, rather than a factor associated with the attack success rate. Furthermore, our experiments reveal a bias momentum effect, where agents resist explicit steering toward attack families that conflict with their bias. This forced distribution shift does not yield measurable improvements in attack performance. To ensure reproducibility and facilitate future research, we release an interactive result dashboard at https://trustworthyai.co.kr/CyBiasBench/ and a reproducibility artifact with aggregated session-level statistics and full evaluation scripts at https://github.com/Harry24k/CyBiasBench.

📄 PDF Abstract BibTeX arXiv:2605.07830

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks

2024-10-20 · Daniel Ayzenshteyn, Roy Weiss, Yisroel Mirsky

As large language models (LLMs) continue to evolve, their potential use in automating cyberattacks becomes increasingly likely. With capabilities such as reconnaissance, exploitation, and command execution, LLMs could so…

Towards a Multi-Agent Simulation of Cyber-attackers and Cyber-defenders Battles

2025-06-05 · Julien Soulé, Jean-Paul Jamont, Michel Occello, Paul Théron 외

As cyber-attacks show to be more and more complex and coordinated, cyber-defenders strategy through multi-agent approaches could be key to tackle against cyber-attacks as close as entry points in a networked system. This…

Learning to Defend by Attacking (and Vice-Versa): Transfer of Learning in Cybersecurity Games

2023-06-03 · Tyler Malloy, Cleotilde Gonzalez

Designing cyber defense systems to account for cognitive biases in human decision making has demonstrated significant success in improving performance against human attackers. However, much of the attention in this area …

Decision MakingLearning Theory

AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges

2026-06-12 · Fengyu Liu, Jiarun Dai, Yihe Fan, Wuyuao Mai 외 arxiv

Frontier AI systems are increasingly capable of cybersecurity tasks, including codebase inspection, vulnerability detection, and exploitation. However, evaluating their offensive capabilities remains constrained by limit…

Vulnerability Detection

When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs

2024-10-18 · Hanna Kim, Minkyoo Song, Seung Ho Na, Seungwon Shin 외

Recent advancements in Large Language Models (LLMs) have established them as agentic systems capable of planning and interacting with various tools. These LLM agents are often paired with web-based tools, enabling access…