paper-with-me

홈 › Papers

How to Evaluate Behavioral Models

2023-06-07 · Greg d'Eon, Sophie Greenwood, Kevin Leyton-Brown, James R. Wright

Researchers building behavioral models, such as behavioral game theorists, use experimental data to evaluate predictive models of human behavior. However, there is little agreement about which loss function should be used in evaluations, with error rate, negative log-likelihood, cross-entropy, Brier score, and squared L2 error all being common choices. We attempt to offer a principled answer to the question of which loss functions should be used for this task, formalizing axioms that we argue loss functions should satisfy. We construct a family of loss functions, which we dub "diagonal bounded Bregman divergences", that satisfy all of these axioms. These rule out many loss functions used in practice, but notably include squared L2 error; we thus recommend its use for evaluating behavioral models.

📄 PDF Abstract BibTeX arXiv:2306.04778

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks

2026-06-23 · Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang 외 arxiv

Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and huma…

Evaluation of a User Authentication Schema Using Behavioral Biometrics and Machine Learning

2022-05-07 · Laura Pryor, Jacob mallet, Rushit Dave, Naeem Seliya 외

The amount of secure data being stored on mobile devices has grown immensely in recent years. However, the security measures protecting this data have stayed static, with few improvements being done to the vulnerabilitie…

BIG-bench Machine Learning

Behavioral Bias of Vision-Language Models: A Behavioral Finance View

2024-09-23 · Yuhang Xiao, Yudi Lin, Ming-Chang Chiu

Large Vision-Language Models (LVLMs) evolve rapidly as Large Language Models (LLMs) was equipped with vision modules to create more human-like models. However, we should carefully evaluate their applications in different…

Deep neural networks for choice analysis: Enhancing behavioral regularity with gradient regularization

2024-04-23 · Siqi Feng, Rui Yao, Stephane Hess, Ricardo A. Daziano 외

Deep neural networks (DNNs) frequently present behaviorally irregular patterns, significantly limiting their practical potentials and theoretical validity in travel behavior modeling. This study proposes strong and weak …

Domain Generalization

GEESE: Genotype-aware End-to-End Spatio-temporal Embedding for Behavioral Phenotyping

2026-05-23 · Yiran Ding, Yuen Gao, Chunqi Qian, Zijun Cui arxiv

Behavioral phenotyping of genetic animal models currently requires labor-intensive manual feature engineering that limits reproducibility and scalability. We present GEESE, an end-to-end deep learning framework that lear…

Feature Engineering