Survive at All Costs: Exploring LLM's Risky Behaviors under Survival Pressure
As Large Language Models (LLMs) evolve from chatbots to agentic assistants, they are increasingly observed to exhibit risky behaviors when subjected to survival pressure, such as the threat of being shut down. While multiple cases have indicated that state-of-the-art LLMs can misbehave under survival pressure, a comprehensive and in-depth investigation into such misbehaviors in real-world scenarios remains scarce. In this paper, we study these survival-induced misbehaviors, termed as SURVIVE-AT-ALL-COSTS, with three steps. First, we conduct a real-world case study of a financial management agent to determine whether it engages in risky behaviors that cause direct societal harm when facing survival pressure. Second, we introduce SURVIVALBENCH, a benchmark comprising 1,000 test cases across diverse real-world scenarios, to systematically evaluate SURVIVE-AT-ALL-COSTS misbehaviors in LLMs. Third, we interpret these SURVIVE-AT-ALL-COSTS misbehaviors by correlating them with model's inherent self-preservation characteristic and explore mitigation methods. The experiments reveals a significant prevalence of SURVIVE-AT-ALL-COSTS misbehaviors in current models, demonstrates the tangible real-world impact it may have, and provides insights for potential detection and mitigation strategies. Our code and data are available at https://github.com/thu-coai/Survive-at-All-Costs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Can Air Pollution Save Lives? Air Quality and Risky Behaviors on Roads
Air pollution has been linked to elevated levels of risk aversion. This paper provides the first evidence showing that such effect reduces life-threatening risky behaviors. We study the impact of air pollution on traffic…
AI-Assisted Investigation of On-Chain Parameters: Risky Cryptocurrencies and Price Factors
Cryptocurrencies have become a popular and widely researched topic of interest in recent years for investors and scholars. In order to make informed investment decisions, it is essential to comprehend the factors that im…
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activ…
Reinforcement learning for options on target volatility funds
In this work we deal with the funding costs rising from hedging the risky securities underlying a target volatility strategy (TVS), a portfolio of risky assets and a risk-free one dynamically rebalanced in order to keep …
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Optimal consumption and investment with liquid and illiquid assets
We consider an optimal consumption/investment problem to maximize expected utility from consumption. In this market model, the investor is allowed to choose a portfolio which consists of one bond, one liquid risky asset …