paper-with-me

홈 › Papers

Beyond Averages: Evaluating LLMs on Human Survey Replication at the Distributional Level

2026-06-08 · Jeonghyeon Moon, Jiwon Kim, Yeheum Lah, Yoonju Han, Yuncheol Kang arxiv

LLMs are increasingly used to simulate human survey responses, but prior work has mainly evaluated replication using mean-level or aggregate agreement, offering limited insight into whether LLMs reproduce the variability of human behavior. We evaluate LLM-based survey replication at the distributional level using a non-public 2010 consumer choice experiment on Korean instant noodle purchases, a setting unlikely to overlap with model training data. We evaluate three response variables of differing statistical type: binary purchase incidence, categorical brand choice, and count purchase quantity. For each, we compare human and LLM responses at mean-level, pattern, and distributional alignment, and against reference baselines from the human data alone. LLMs reproduce condition-level patterns reasonably well but fail to capture distributional structure: for purchase quantity, no model beats a condition-insensitive baseline that simply matches the pooled human distribution. Because models that match human means well can still produce distributions further from humans than this baseline, mean-based evaluation alone can be actively misleading. Replication also varies with input configuration, with structured personas and multimodal inputs improving alignment while explicit reasoning prompting degrades it monotonically.

📄 PDF Abstract BibTeX arXiv:2606.09013

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

2024-04-02 · Philipp Mondorf, Barbara Plank

Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite …

Survey

Beyond Marginal Distributions: A Framework to Evaluate the Representativeness of Demographic-Aligned LLMs

2026-01-22 · Tristan Williams, Franziska Weeber, Sebastian Padó, Alan Akbik arxiv

Large language models are increasingly used to represent human opinions, values, or beliefs, and their steerability towards these ideals is an active area of research. Existing work focuses predominantly on aligning marg…

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

2026-01-13 · Guo-Biao Zhang, Ding-Yuan Liu, Da-Yi Wu, Tian Lan 외 arxiv

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first const…

Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare

2025-04-11 · Yonchanok Khaokaew, Flora D. Salim, Andreas Züfle, Hao Xue 외

Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or…

Decision MakingPrompt EngineeringSurvey

Aligning Large Language Models with Human: A Survey

2023-07-24 · YuFei Wang, Wanjun Zhong, Liangyou Li, Fei Mi 외

Large Language Models (LLMs) trained on extensive textual corpora have emerged as leading solutions for a broad array of Natural Language Processing (NLP) tasks. Despite their notable performance, these models are prone …

Survey