paper-with-me

Papers

A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains

2025-08-18 · Xianren Zhang, Shreyas Prasad, Di Wang, Qiuhai Zeng, Suhang Wang, Wenbo Yan, Mat Hans arxiv

Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been introduced. However, current benchmarks in the e-commerce domain face two major problems. First, they primarily focus on product search tasks (e.g., Find an Apple Watch), failing to capture the broader range of functionalities offered by real-world e-commerce platforms such as Amazon, including account management and gift card operations. Second, existing benchmarks typically evaluate whether the agent completes the user query, but ignore the potential risks involved. In practice, web agents can make unintended changes that negatively impact the user account or status. For instance, an agent might purchase the wrong item, delete a saved address, or incorrectly configure an auto-reload setting. To address these gaps, we propose a new benchmark called Amazon-Bench. To generate user queries that cover a broad range of tasks, we propose a data generation pipeline that leverages webpage content and interactive elements (e.g., buttons, check boxes) to create diverse, functionality-grounded user queries covering tasks such as address management, wish list management, and brand store following. To improve the agent evaluation, we propose an automated evaluation framework that assesses both the performance and the safety of web agents. We systematically evaluate different agents, finding that current agents struggle with complex queries and pose safety risks. These results highlight the need for developing more robust and reliable web agents.

📄 PDF Abstract BibTeX arXiv:2508.15832

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce

2026-02-01 · Alberto Castelo, Zahra Zanjani Foumani, Ailin Fan, Keat Yang Koay 외 arxiv

A/B testing remains the gold standard for evaluating e-commerce UI changes, yet it diverts traffic, takes weeks to reach significance, and risks harming user experience. We introduce SimGym, a scalable system for rapid o…

Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules

2025-09-28 · Chenyu Zhou, Xiaoming Shi, Hui Qiu, Xiawu Zheng 외 arxiv

E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are introduced for evaluating LLM agents in…

Evaluating Multi-Turn Bargain Skills in LLM-Based Seller Agent

2025-09-08 · Issue Yishu Wang, Kakam Chong, Xiaofeng Wang, Xu Yan 외 arxiv

In online second-hand marketplaces, multi-turn bargaining is a crucial part of seller-buyer interactions. Large Language Models (LLMs) can act as seller agents, negotiating with buyers on behalf of sellers under given bu…

Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks

2026-03-16 · Zijian Yu, Kejun Xiao, Huaipeng Zhao, Tao Luo 외 arxiv

In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing user preferences from long-horizon conversations is critical. However, pr…

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

2026-05-19 · Han Li, Vibhor Malik, Zahra Zanjani Foumani, Alberto Castelo 외 arxiv

A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistical significance, and risks degrading user experience. We present SimG…