paper-with-me

홈 › Papers

OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

2025-06-05 · Ziyi Wang, Yuxuan Lu, Wenbo Li, Amirali Amini, Bo Sun, Yakov Bart, Weimin Lyu, Jiri Gesi, Tian Wang, Jing Huang, Yu Su, Upol Ehsan, Malihe Alikhani, Toby Jia-Jun Li, Lydia Chilton, Dakuo Wang

Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generating ``believable'' human behaviors, evaluating their ability to mimic real user behaviors remains an open challenge, largely due to the lack of high-quality, publicly available datasets that capture both the observable actions and the internal reasoning of an actual human user. To address this gap, we introduce OPERA, a novel dataset of Observation, Persona, Rationale, and Action collected from real human participants during online shopping sessions. OPERA is the first public dataset that comprehensively captures: user personas, browser observations, fine-grained web actions, and self-reported just-in-time rationales. We developed both an online questionnaire and a custom browser plugin to gather this dataset with high fidelity. Using OPERA, we establish the first benchmark to evaluate how well current LLMs can predict a specific user's next action and rationale with a given persona and <observation, action, rationale> history. This dataset lays the groundwork for future research into LLM agents that aim to act as personalized digital twins for human.

📄 PDF Abstract BibTeX arXiv:2506.05606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping

2025-10-08 · Ziyi Wang, Yuxuan Lu, Yimeng Zhang, Jing Huang 외 arxiv

Simulating step-wise human behavior with Large Language Models (LLMs) has become an emerging research direction, enabling applications in various practical domains. While prior methods, including prompting, supervised fi…

Reinforcement Learning

Improving Language Model Personas via Rationalization with Psychological Scaffolds

2025-04-25 · Brihi Joshi, Xiang Ren, Swabha Swayamdipta, Rik Koncel-Kedziorski 외

Language models prompted with a user description or persona are being used to predict the user's preferences and opinions. However, existing approaches to building personas mostly rely on a user's demographic attributes …

Language ModelingLanguage Modelling

Persona Prompting as a Lens on LLM Social Reasoning

2026-01-28 · Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov, Evelyn Luise Brinkmann 외 arxiv

For socially sensitive tasks like hate speech detection, the quality of explanations from Large Language Models (LLMs) is crucial for factors like user trust and model alignment. While Persona prompting (PP) is increasin…

Hate Speech Detection

Enhancing the Rationale-Input Alignment for Self-explaining Rationalization

2023-12-07 · Wei Liu, Haozhao Wang, Jun Wang, Zhiying Deng 외

Rationalization empowers deep learning models with self-explaining capabilities through a cooperative game, where a generator selects a semantically consistent subset of the input as a rationale, and a subsequent predict…

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

2026-05-28 · Rahul Bissa, Abhishek Vyas, Yash Jain arxiv

We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Action, Rationale), a 12,929-tuple corpus of screen-anchored behavioural…