paper-with-me

홈 › Papers

Out-of-Distribution Generalization of Risk Aversion in Language Models

2026-07-02 · Kristina Zhang, Junior Chinomso Okoroafor, Benjamin Maltbie, Andrew Lin, Abhitej Bokka, Elliott Thornley arxiv

Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the downsides of any misalignment. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring how well risk aversion generalizes out of distribution. We then offer some initial results. Using a variety of methods to make Qwen3-8B choose risk-aversely when the stakes are low, we find that we can induce substantial risk aversion when the stakes are astronomically high. Our models' learned risk aversion generalizes at least partially across 98 orders of magnitude. From a baseline 2% rate of choosing a safe `Cooperate' option, we see rates around 70% (SFT and tie training), 52% (DPO), and 39% (activation steering). In another experiment, our fine-tuned reward model reliably scores risk-averse reasoning above risk-neutral or excessively risk-averse alternatives (99.6% pairwise accuracy). We replicate these effects at different scales (Qwen3-1.7B and Qwen3-14B) and across model families (Gemma-3-12B-IT and Llama-3.1-8B-Instruct). Overall, we find that risk aversion learned at low stakes can generalize OOD to astronomically high stakes, though not yet consistently enough to serve as a reliable failsafe. Achieving that level of consistency is an open problem.

📄 PDF Abstract BibTeX arXiv:2607.02755

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inequality and risk aversion in economies open to altruistic attitudes

2016-05-11

This paper attempts to find a relationship between agents' risk aversion and inequality of incomes. Specifically, a model is proposed for the evolution in time of surplus/deficit distribution, and the long-time distribut…

Evolutionary Foundation for Heterogeneity in Risk Aversion

2021-10-21 · Yuval Heller, Ilan Nehama

We examine the evolutionary basis for risk aversion with respect to aggregate risk. We study populations in which agents face choices between alternatives with different levels of aggregate risk. We show that the choices…

Procurements with Bidder Asymmetry in Cost and Risk-Aversion

2021-11-08 · Gaurab Aryal, Hanna Charankevich, Seungwon Jeong, Dong-Hyuk Kim

We propose an empirical method to analyze data from first-price procurements where bidders are asymmetric in their risk-aversion (CRRA) coefficients and distributions of private costs. Our Bayesian approach evaluates the…

counterfactualData Augmentation

Eliciting Risk Aversion with Inverse Reinforcement Learning via Interactive Questioning

2023-08-16 · Ziteng Cheng, Anthony Coache, Sebastian Jaimungal

This paper proposes a novel framework for identifying an agent's risk aversion using interactive questioning. Our study is conducted in two scenarios: a one-period case and an infinite horizon case. In the one-period cas…

reinforcement-learningReinforcement Learning

Distributionally Robust Optimization: A Review

2019-08-13 · Hamed Rahimian, Sanjay Mehrotra

The concepts of risk-aversion, chance-constrained optimization, and robust optimization have developed significantly over the last decade. Statistical learning community has also witnessed a rapid theoretical and applied…