paper-with-me

Papers

WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks

2026-04-07 · Guruprasad Viswanathan Ramesh, Asmit Nayak, Basieem Siddique, Kassem Fawaz arxiv

Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate general-purpose performance~(e.g., WebArena) or safety against malicious actions~(e.g., SafeArena), no existing framework assesses an agent's ability to successfully execute user-facing website security and privacy tasks, such as managing cookie preferences, configuring privacy-sensitive account settings, or revoking inactive sessions. To address this gap, we introduce WebSP-Eval, an evaluation framework for measuring web agent performance on website security and privacy tasks. WebSP-Eval comprises 1) a manually crafted task dataset of 200 task instances across 28 websites; 2) a robust agentic system supporting account and initial state management across runs using a custom Google Chrome extension; and 3) an automated evaluator. We evaluate a total of 8 web agent instantiations using state-of-the-art multimodal large language models, conducting a fine-grained analysis across websites, task categories, and UI elements. Our evaluation reveals that current models suffer from limited autonomous exploration capabilities to reliably solve website security and privacy tasks, and struggle with specific task categories and websites. Crucially, we identify stateful UI elements are a primary reason for agent failure, with toggles causing more than 45% task failure across many models.

📄 PDF Abstract BibTeX arXiv:2604.06367

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WebSpline: Structure-Informed Splines for Real-Time 3D Gaussians from Monocular Videos

2026-06-01 · Jongmin Park, Jeonghwan Yun, Minh-Quan Viet Bui, Munchurl Kim arxiv

Dynamic scene reconstruction from monocular videos remains highly challenging, as existing methods often struggle to balance global structural coherence and local fine-grained details under limited multi-view cues. To ad…

WebSplatter: Enabling Cross-Device Efficient Gaussian Splatting in Web Browsers via WebGPU

2026-02-03 · Yudong Han, Chao Xu, Xiaodan Ye, Weichen Bi 외 arxiv

We present WebSplatter, an end-to-end GPU rendering pipeline for the heterogeneous web ecosystem. Unlike naive ports, WebSplatter introduces a wait-free hierarchical radix sort that circumvents the lack of global atomics…

MalURLBench: A Benchmark Evaluating Agents' Vulnerabilities When Processing Web URLs

2026-01-26 · Dezhang Kong, Zhuxi Wu, Shiqi Liu, Zhicheng Tan 외 arxiv

LLM-based web agents have become increasingly popular for their utility in daily life and work. However, they exhibit critical vulnerabilities when processing malicious URLs: accepting a disguised malicious URL enables s…

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

2024-06-19 · Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner 외

AI agents aim to solve complex tasks by combining text-based reasoning with external tool calls. Unfortunately, AI agents are vulnerable to prompt injection attacks where data returned by external tools hijacks the agent…

Mind2Web: Towards a Generalist Agent for the Web

2023-06-09 · NeurIPS 2023 11 · Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen 외

We introduce Mind2Web, the first dataset for developing and evaluating generalist agents for the web that can follow language instructions to complete complex tasks on any website. Existing datasets for web agents either…