paper-with-me

홈 › Papers

Emergence WebVoyager: Toward Consistent and Transparent Evaluation of (Web) Agents in The Wild

2026-03-30 · Deepak Akkil, Mowafak Allaham, Amal Raj, Tamer Abuelsaad, Ravi Kokku arxiv

Reliable evaluation of AI agents operating in complex, real-world environments requires methodologies that are robust, transparent, and contextually aligned with the tasks agents are intended to perform. This study identifies persistent shortcomings in existing AI agent evaluation practices that are particularly acute in web agent evaluation, as exemplified by our audit of WebVoyager, including task-framing ambiguity and operational variability that hinder meaningful and reproducible performance comparisons. To address these challenges, we introduce Emergence WebVoyager, an enhanced version of the WebVoyager benchmark that standardizes evaluation methodology through clear guidelines for task instantiation, failure handling, annotation, and reporting. Emergence WebVoyager achieves an inter-annotator agreement of 95.9\%, indicating improved clarity and reliability in both task formulation and evaluation. Applying this framework to evaluate OpenAI Operator reveals substantial performance variation across domains and task types, with an overall success rate of 68.6\%, substantially lower than the 87\% previously reported by OpenAI, demonstrating the utility of our approach for more rigorous and comparable web agent evaluation.

📄 PDF Abstract BibTeX arXiv:2603.29020

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

2024-01-25 · Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu 외

The rapid advancement of large language models (LLMs) has led to a new era marked by the development of autonomous applications in real-world scenarios, which drives innovation in creating advanced web agents. Existing w…

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

2026-04-09 · Tanmay Gupta, Piper Wolters, Zixian Ma, Peter Sushko 외 arxiv

Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, the most capable web agents today rely on…

Referring ExpressionQuestion Answering

Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems

2024-07-17 · Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan 외

AI Agents are changing the way work gets done, both in consumer and enterprise domains. However, the design patterns and architectures to build highly capable agents or multi-agent systems are still developing, and the u…

Autonomous Web NavigationDenoising

AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines

2026-02-15 · Yifan Wu, Yiran Peng, Yiyu Chen, Jianhao Ruan 외 arxiv

The performance of autonomous Web GUI agents heavily relies on the quality and quantity of their training data. However, a fundamental bottleneck persists: collecting interaction trajectories from real-world websites is …

Visual Test-time Scaling for GUI Agent Grounding

2025-05-01 · Tiange Luo, Lajanugen Logeswaran, Justin Johnson, Honglak Lee

We introduce RegionFocus, a visual test-time scaling approach for Vision Language Model Agents. Understanding webpages is challenging due to the visual complexity of GUI images and the large number of interface elements,…

Language ModelingLanguage Modelling