paper-with-me

홈 › Papers

StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability

2026-03-27 · Haoyue Bai, Dong Wang, Long Chen, Bingguang Hao, Pengyang Shao, Yonghui Yang, Yicheng He, Chenyi Zhuang arxiv

Large language model-based web agents have demonstrated strong performance on realistic web interaction tasks. However, existing evaluations are predominantly conducted under relatively stable and well-behaved interaction conditions, which may overestimate agent robustness. High task success in such idealized settings does not necessarily reflect performance under realistic web interaction. To address this limitation, we introduce a diagnostic stress-testing benchmark for web agents. We first construct realistic and controllable web environments that provide clean and stable interaction workflows as reference baselines. We then introduce structured and controlled perturbations that emulate interaction variability, including shifting layouts, altered interaction semantics, and execution disruptions. By comparing agent behavior between clean and perturbed settings, our framework enables systematic diagnosis of robustness under what-if interaction scenarios. Through extensive evaluation of state-of-the-art multimodal web agents, we show that stress-based evaluation exposes failure modes and substantial robustness gaps that remain hidden under clean benchmark conditions.

📄 PDF Abstract BibTeX arXiv:2604.16385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

2026-06-03 · Haoyu Sun, Wenxuan Wang, Mingyang Song, Jujie He 외 arxiv

Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeasible. Yet existing agent evaluations often report only end-to-end su…

Back to Basics: Revisiting ASR in the Age of Voice Agents

2026-03-26 · Geeyang Tay, Wentao Ma, Jaewon Lee, Yuzhi Tang 외 arxiv

Automatic speech recognition (ASR) systems have achieved near-human accuracy on curated benchmarks, yet still fail in real-world voice agents under conditions that current evaluations do not systematically cover. Without…

Speech Recognition

Embeddings to Diagnosis: Latent Fragility under Agentic Perturbations in Clinical LLMs

2025-07-27 · Raj Krishnan Vijayaraj arxiv

LLMs for clinical decision support often fail under small but clinically meaningful input shifts such as masking a symptom or negating a finding, despite high performance on static benchmarks. These reasoning failures fr…

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

2026-05-13 · Yifei Zhu arxiv

Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into open-ended exploration. Yet real world use requires models to discover…

Information RetrievalQuestion Answering

MedExpMem: Adapting Experience Memory for Differential Diagnosis

2026-05-20 · Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Yannian Gu 외 arxiv

Experienced physicians develop diagnostic expertise through clinical practice, acquiring not only disease knowledge but also the ability to differentiate confusable conditions. Current medical vision-language models (VLM…