paper-with-me

Papers

WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making

2026-03-22 · Zongjie Li, Chaozheng Wang, Yuchong Xie, Pingchuan Ma, Shuai Wang arxiv

Large Language Models are increasingly being considered for deployment in safety-critical military applications. However, current benchmarks suffer from structural blindspots that systematically overestimate model capabilities in real-world tactical scenarios. Existing frameworks typically ignore strict legal constraints based on International Humanitarian Law (IHL), omit edge computing limitations, lack robustness testing for fog of war, and inadequately evaluate explicit reasoning. To address these vulnerabilities, we present WARBENCH, a comprehensive evaluation framework establishing a foundational tactical baseline alongside four distinct stress testing dimensions. Through a large scale empirical evaluation of nine leading models on 136 high-fidelity historical scenarios, we reveal severe structural flaws. First, baseline tactical reasoning systematically collapses under complex terrain and high force asymmetry. Second, while state of the art closed source models maintain functional compliance, edge-optimized small models expose extreme operational risks with legal violation rates approaching 70 percent. Furthermore, models experience catastrophic performance degradation under 4-bit quantization and systematic information loss. Conversely, explicit reasoning mechanisms serve as highly effective structural safeguards against inadvertent violations. Ultimately, these findings demonstrate that current models remain fundamentally unready for autonomous deployment in high stakes tactical environments.

📄 PDF Abstract BibTeX arXiv:2603.21280

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

2026-04-30 · Sydney Johns, Heng Jin, Chaoyu Zhang, Y. Thomas Hou 외 arxiv

Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hold significant potential to enhance decision making, coordination, an…

Decision Making

Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making

2025-10-03 · Toby Drinkall arxiv

As military organisations consider integrating large language models (LLMs) into command and control (C2) systems for planning and decision support, understanding their behavioural tendencies is critical. This study deve…

Measuring and Eliminating Refusals in Military Large Language Models

2026-02-18 · Jack FitzGerald, Dylan Bates, Aristotelis Lazaridis, Aman Sharma 외 arxiv

Military Large Language Models (LLMs) must provide accurate information to the warfighter in time-critical and dangerous situations. However, today's LLMs are imbued with safety behaviors that cause the LLM to refuse man…

Neuro-Symbolic AI for Military Applications

2024-08-17 · Desta Haileselassie Hagos, Danda B. Rawat

Artificial Intelligence (AI) plays a significant role in enhancing the capabilities of defense systems, revolutionizing strategic decision-making, and shaping the future landscape of military operations. Neuro-Symbolic A…

Decision Making

Balancing Power and Ethics: A Framework for Addressing Human Rights Concerns in Military AI

2024-11-10 · Mst Rafia Islam, Azmine Toushik Wasi

AI has made significant strides recently, leading to various applications in both civilian and military sectors. The military sees AI as a solution for developing more effective and faster technologies. While AI offers b…

EthicsHumanitarian