paper-with-me

홈 › Papers

Managing Escalation in Off-the-Shelf Large Language Models

2025-08-01 · Sebastian Elbaum, Jonathan Panter arxiv

U.S. national security customers have begun to utilize large language models, including enterprise versions of ``off-the-shelf'' models (e.g., ChatGPT) familiar to the public. This uptake will likely accelerate. However, recent studies suggest that off-the-shelf large language models frequently suggest escalatory actions when prompted with geopolitical or strategic scenarios. We demonstrate two simple, non-technical interventions to control these tendencies. Introducing these interventions into the experimental wargame design of a recent study, we substantially reduce escalation throughout the game. Calls to restrict the use of large language models in national security applications are thus premature. The U.S. government is already, and will continue, employing large language models for scenario planning and suggesting courses of action. Rather than warning against such applications, this study acknowledges the imminent adoption of large language models, and provides actionable measures to align them with national security goals, including escalation management.

📄 PDF Abstract BibTeX arXiv:2508.01056

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Escalation Prediction using Feature Engineering: Addressing Support Ticket Escalations within IBM's Ecosystem

2020-10-12 · Lloyd Montgomery

Large software organizations handle many customer support issues every day in the form of bug reports, feature requests, and general misunderstandings as submitted by customers. Strategies to gather, analyze, and negotia…

BIG-bench Machine LearningFeature EngineeringManagement

Customer Support Ticket Escalation Prediction using Feature Engineering

2020-10-10 · Lloyd Montgomery, Daniela Damian, Tyson Bulmer, Shaikh Quader

Understanding and keeping the customer happy is a central tenet of requirements engineering. Strategies to gather, analyze, and negotiate requirements are complemented by efforts to manage customer input after products h…

Feature EngineeringManagementPrediction

Escalation Risks from Language Models in Military and Diplomatic Decision-Making

2024-01-07 · Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth 외

Governments are increasingly considering integrating autonomous AI agents in high-stakes military and foreign-policy decision-making, especially with the emergence of advanced generative AI models like GPT-4. Our work ai…

Decision MakingLanguage Modelling

Emotionally-Aware Agents for Dispute Resolution

2025-08-28 · Sushrita Rakshit, James Hale, Kushal Chawla, Jeanne M. Brett 외 arxiv

In conflict, people use emotional expressions to shape their counterparts' thoughts, feelings, and actions. This paper explores whether automatic text emotion recognition offers insight into this influence in the context…

Emotion Recognition

Modeling Clinical Concern Trajectories in Language Model Agents

2026-04-30 · Sukesh Subaharan, Venkatesan VS, Murugadasan P, Sivakumar D 외 arxiv

Large language model (LLM) agents deployed in clinical settings often exhibit abrupt, threshold-driven behavior, offering little visibility into accumulating risk prior to escalation. In real-world care, however, clinici…