paper-with-me

홈 › Papers

AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping

2026-02-12 · Sunghwan Kim, Ryang Heo, Yongsik Seo, Jinyoung Yeo, Dongha Lee arxiv

The proliferation of e-commerce has made web shopping platforms key gateways for customers navigating the vast digital marketplace. Yet this rapid expansion has led to a noisy and fragmented information environment, increasing cognitive burden as shoppers explore and purchase products online. With promising potential to alleviate this challenge, agentic systems have garnered growing attention for automating user-side tasks in web shopping. Despite significant advancements, existing benchmarks fail to comprehensively evaluate how well agentic systems can curate products in open-web settings. Specifically, they have limited coverage of shopping scenarios, focusing only on simplified single-platform lookups rather than exploratory search. Moreover, they overlook personalization in evaluation, leaving unclear whether agents can adapt to diverse user preferences in realistic shopping contexts. To address this gap, we present AgenticShop, the first benchmark for evaluating agentic systems on personalized product curation in open-web environment. Crucially, our approach features realistic shopping scenarios, diverse user profiles, and a verifiable, checklist-driven personalization evaluation framework. Through extensive experiments, we demonstrate that current agentic systems remain largely insufficient, emphasizing the need for user-side systems that effectively curate tailored products across the modern web.

📄 PDF Abstract BibTeX arXiv:2602.12315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Silent Failures in Multi-Agentic AI Trajectories

2025-11-06 · Divya Pathak, Harshit Kumar, Anuska Roy, Felix George 외 arxiv

Multi-Agentic AI systems, powered by large language models (LLMs), are inherently non-deterministic and prone to silent failures such as drift, cycles, and missing details in outputs, which are difficult to detect. We in…

Anomaly Detection

APeB: Benchmarking Personalization Ability of Large Language Model Agents

2026-07-03 · Garry Yang, Zizhe Chen, Xinru Chen, Yongqiang Chen 외 arxiv

LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract preferences from noisy interaction histories, and select among comp…

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

2026-04-02 · Smriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali 외 arxiv

Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and fidelity: online A/B testing takes weeks and risks user experience, shadow deplo…

CurateEvo: Data-Curation Evolving for Agentic Post-Training

2026-07-07 · Dingzirui Wang, Xuanliang Zhang, Keyan Xu, Qingfu Zhu 외 arxiv

Large language model (LLM) agents require post-training methods that can improve long-horizon decision making from environment feedback. However, existing agentic post-training pipelines often treat data curation as a fi…

Reinforcement LearningData AugmentationDecision Making

AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems

2025-05-26 · Yu Shang, Peijie Liu, Yuwei Yan, Zijing Wu 외

The emergence of agentic recommender systems powered by Large Language Models (LLMs) represents a paradigm shift in personalized recommendations, leveraging LLMs' advanced reasoning and role-playing capabilities to enabl…

BenchmarkingRecommendation Systems