Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making
Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human decision-making's variability and adaptability. We propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, and Imitation) to examine how LLM agents adapt under different levels of external guidance and human-derived noise. We validate the framework on two classic economics tasks, irrationality in the second-price auction and decision bias in the newsvendor problem, showing behavioral gaps between LLMs and humans. We find that LLMs, by default, converge on stable and conservative strategies that diverge from observed human behaviors. Risk-framed instructions impact LLM behavior predictably but do not replicate human-like diversity. Incorporating human data through in-context learning narrows the gap but fails to reach human subjects' strategic variability. These results highlight a persistent alignment gap in behavioral fidelity and suggest that future LLM evaluations should consider more process-level realism. We present a process-oriented approach for assessing LLMs in dynamic decision-making tasks, offering guidance for their application in synthetic data for social science research.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A Novel Composite Resilience Indicator for Decentralized Infrastructure Systems (CRI-DS)
Resilience is a key driver for planning adaptation strategies to mitigate risks due to both natural and anthropogenic hazards. The effectiveness of a resilience-driven decision-making strategy for adapting systems agains…
Decision MakingLightweight Adaptation for LLM-based Technical Service Agent: Latent Logic Augmentation and Robust Noise Reduction
Adapting Large Language Models in complex technical service domains is constrained by the absence of explicit cognitive chains in human demonstrations and the inherent ambiguity arising from the diversity of valid respon…
Computational EfficiencyReinforcement LearningTrajectory ModelingSpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras
RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB-to-RAW conver…
Object DetectionX2HDR: HDR Image Generation in a Perceptually Uniform Space
High-dynamic-range (HDR) formats and displays are becoming increasingly prevalent, yet state-of-the-art image generators (e.g., Stable Diffusion and FLUX) typically remain limited to low-dynamic-range (LDR) output due to…
Image GenerationPrior Mismatch and Adaptation in PnP-ADMM with a Nonconvex Convergence Analysis
Plug-and-Play (PnP) priors is a widely-used family of methods for solving imaging inverse problems by integrating physical measurement models with image priors specified using image denoisers. PnP methods have been shown…
Domain AdaptationImage Super-ResolutionSuper-Resolution