Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search Systems
LLM-driven GUI agents are increasingly used in production systems to automate workflows and simulate users for evaluation and optimization. Yet most GUI-agent evaluations emphasize task success and provide limited evidence on whether agents interact in human-like ways. We present a trace-level evaluation framework that compares human and agent behavior across (i) task outcome and effort, (ii) query formulation, and (iii) navigation across interface states. We instantiate the framework in a controlled study in a production audio-streaming search application, where 39 participants and a state-of-the-art GUI agent perform ten multi-hop search tasks. The agent achieves task success comparable to participants and generates broadly aligned queries, but follows systematically different navigation strategies: participants exhibit content-centric, exploratory behavior, while the agent is more search-centric and low-branching. These results show that outcome and query alignment do not imply behavioral alignment, motivating trace-level diagnostics when deploying GUI agents as proxies for users in production search systems.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Patient Similarity Analysis with Longitudinal Health Data
Healthcare professionals have long envisioned using the enormous processing powers of computers to discover new facts and medical knowledge locked inside electronic health records. These vast medical archives contain tim…
Decision MakingUnderstanding the Progression of Educational Topics via Semantic Matching
Education systems are dynamically changing to accommodate technological advances, industrial and societal needs, and to enhance students' learning journeys. Curriculum specialists and educators constantly revise taught s…
MathTRACE: Transformer-based user Representations from Attributed Clickstream Event sequences
For users navigating travel e-commerce websites, the process of researching products and making a purchase often results in intricate browsing patterns that span numerous sessions over an extended period of time. The res…
Multi-Task LearningRecommendation SystemsAdaptive User Journeys in Pharma E-Commerce with Reinforcement Learning: Insights from SwipeRx
This paper introduces a reinforcement learning (RL) platform that enhances end-to-end user journeys in healthcare digital tools through personalization. We explore a case study with SwipeRx, the most popular all-in-one a…
ManagementReinforcement Learning (RL)Predicting Children's Travel Modes for School Journeys in Switzerland: A Machine Learning Approach Using National Census Data
Children's travel behavior plays a critical role in shaping long-term mobility habits and public health outcomes. Despite growing global interest, little is known about the factors influencing travel mode choice of child…