paper-with-me

홈 › Papers

Measuring Behavior Portability in Large Language Models

2026-06-22 · Tianjia Dong, Nadav Kunievsky, James A. Evans arxiv

Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface presentation. This sensitivity renders suite-based evaluation fragile and raises a fundamental question of behavioral portability: how well does a behavioral mapping learned in one decision environment informative on another that preserves the same underlying incentive structure? We introduce a formal framework to measure this property. Our protocol fits an interpretable behavioral model on data pooled from a set of source environments and evaluates its out-of-sample predictive performance in a held-out target environment, benchmarking against an oracle trained directly on target data. Portability is quantified via a loss-agnostic measure that delivers worst-case bounds on the performance of the induced prediction-action mapping in the target environment. In controlled experiments spanning seven canonical economic decision problems, we document substantial and systematic portability losses, suggesting that behavioral characterizations of LLMs obtained in one decision environment cannot be assumed to transfer reliably to structurally equivalent alternatives.

📄 PDF Abstract BibTeX arXiv:2606.22797

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-Demographic Portability of Deep NLP-Based Depression Models

2024-12-26 · Tomek Rutowski, Elizabeth Shriberg, Amir Harati, Yang Lu 외

Deep learning models are rapidly gaining interest for real-world applications in behavioral health. An important gap in current literature is how well such models generalize over different populations. We study Natural L…

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

2026-05-17 · Qiuhe Hong, Yuyang Liu, Shuo Yang, Tiantian Peng 외 arxiv

Vision-Language Models in Continual Learning (VLM-CL) aim to continuously adapt to new multimodal tasks while retaining prior knowledge. The emerging paradigm that couples Multimodal Large Language Models (MLLMs) with Re…

Reinforcement LearningContinual Learning

Do Large Language Models Walk Their Talk? Measuring the Gap Between Implicit Associations, Self-Report, and Behavioral Altruism

2025-12-01 · Sandro Andric arxiv

We investigate whether Large Language Models (LLMs) exhibit altruistic tendencies, and critically, whether their implicit associations and self-reports predict actual altruistic behavior. Using a multi-method approach in…

Switching Contexts: Transportability Measures for NLP

2021-05-03 · IWCS (ACL) 2021 6 · Guy Marshall, Mokanarangan Thayaparan, Philip Osborne, Andre Freitas

This paper explores the topic of transportability, as a sub-area of generalisability. By proposing the utilisation of metrics based on well-established statistics, we are able to estimate the change in performance of NLP…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Inference

One Stack, Diverse Vehicles: Checking Safe Portability of Automated Driving Software

2025-01-30 · Vladislav Nenchev

Integrating an automated driving software stack into vehicles with variable configuration is challenging, especially due to different hardware characteristics. Further, to provide software updates to a vehicle fleet in t…

Diversity