paper-with-me

Papers

EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce

2025-12-09 · Rui Min, Zile Qiao, Ze Xu, Jiawen Zhai, Wenyu Gao, Xuanzhong Chen, Haozhen Sun, Zhen Zhang, Xinyu Wang, Hong Zhou, Wenbiao Yin, Bo Zhang, Xuan Zhou, Ming Yan, Yong Jiang, Haicheng Liu, Liang Ding, Ling Zou, Yi R. Fung, Yalong Li, Pengjun Xie arxiv

Foundation agents have rapidly advanced in their ability to reason and interact with real environments, making the evaluation of their core capabilities increasingly important. While many benchmarks have been developed to assess agent performance, most concentrate on academic settings or artificially designed scenarios while overlooking the challenges that arise in real applications. To address this issue, we focus on a highly practical real-world setting, the e-commerce domain, which involves a large volume of diverse user interactions, dynamic market conditions, and tasks directly tied to real decision-making processes. To this end, we introduce EcomBench, a holistic E-commerce Benchmark designed to evaluate agent performance in realistic e-commerce environments. EcomBench is built from genuine user demands embedded in leading global e-commerce ecosystems and is carefully curated and annotated through human experts to ensure clarity, accuracy, and domain relevance. It covers multiple task categories within e-commerce scenarios and defines three difficulty levels that evaluate agents on key capabilities such as deep information retrieval, multi-step reasoning, and cross-source knowledge integration. By grounding evaluation in real e-commerce contexts, EcomBench provides a rigorous and dynamic testbed for measuring the practical capabilities of agents in modern e-commerce.

📄 PDF Abstract BibTeX arXiv:2512.08868

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

2025-10-11 · Zonghao Ying, Yangguang Shao, Jianle Gan, Gan Xu 외 arxiv

Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed in real-world environments, they face serious security risks, motivating the …

RiskWebWorld: A Realistic Interactive Benchmark for GUI Agents in E-commerce Risk Management

2026-04-15 · Renqi Chen, Zeyin Tao, Jianming Guo, Jing Wang 외 arxiv

Graphical User Interface (GUI) agents show strong capabilities for automating web tasks, but existing interactive benchmarks primarily target benign, predictable consumer environments. Their effectiveness in high-stakes,…

Reinforcement Learning

Captions Speak Louder than Images (CASLIE): Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data

2024-10-22 · Xinyi Ling, Bo Peng, Hanwen Du, Zhihui Zhu 외

Leveraging multimodal data to drive breakthroughs in e-commerce applications through Multimodal Foundation Models (MFMs) is gaining increasing attention from the research community. However, there are significant challen…

Domain Adaptation of Foundation LLMs for e-Commerce

2025-01-16 · Christian Herold, Michael Kozielski, Tala Bazazo, Pavel Petrushkov 외

We present the e-Llama models: 8 billion and 70 billion parameter large language models that are adapted towards the e-commerce domain. These models are meant as foundation models with deep knowledge about e-commerce, th…

Domain Adaptation

A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains

2025-08-18 · Xianren Zhang, Shreyas Prasad, Di Wang, Qiuhai Zeng 외 arxiv

Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been introduced. However, current benchmarks in the e-commerce domain face two majo…