paper-with-me

홈 › Papers

How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?

2026-02-06 · Yuxuan Li, Leyang Li, Hao-Ping Lee, Sauvik Das arxiv

A growing body of research assumes that large language model (LLM) agents can serve as proxies for how people form attitudes toward and behave in response to security and privacy (S&P) threats. If correct, these simulations could offer a scalable way to forecast S&P risks in products prior to deployment. We interrogate this assumption using SP-ABCBench, a new benchmark of 30 tests derived from validated S&P human-subject studies, which measures alignment between simulations and human-subjects studies on a 0-100 ascending scale, where higher scores indicate better alignment across three dimensions: Attitude, Behavior, and Coherence. Evaluating twelve LLMs, four persona construction strategies, and two prompting methods, we found that there remains substantial room for improvement: all models score between 50 and 64 on average. Newer, bigger, and smarter models do not reliably do better and sometimes do worse. Some simulation configurations, however, do yield high alignment: e.g., with scores above 95 for some behavior tests when agents are prompted to apply bounded rationality and weigh privacy costs against perceived benefits. We release SP-ABCBench to enable reproducible evaluation as methods improve.

📄 PDF Abstract BibTeX arXiv:2602.18464

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior

2026-05-12 · James Flemings, Murali Annavaram arxiv

Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whet…

Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents

2025-04-24 · Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo 외

The rise of Large Language Models (LLMs) has revolutionized Graphical User Interface (GUI) automation through LLM-powered GUI agents, yet their ability to process sensitive data with limited human oversight raises signif…

LLM Agents Should Employ Security Principles

2025-05-29 · Kaiyuan Zhang, Zian Su, Pin-Yu Chen, Elisa Bertino 외

Large Language Model (LLM) agents show considerable promise for automating complex tasks using contextual reasoning; however, interactions involving multiple agents and the system's susceptibility to prompt injection and…

Large Language Model

EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage

2024-09-17 · Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang 외

Generalist web agents have demonstrated remarkable potential in autonomously completing a wide range of tasks on real websites, significantly boosting human productivity. However, web tasks, such as booking flights, usua…

AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents

2025-03-12 · Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Maya Pavlova 외

LLM-powered AI agents are an emerging frontier with tremendous potential to increase human productivity. However, empowering AI agents to take action on their user's behalf in day-to-day tasks involves giving them access…