paper-with-me

Papers

An Executable Benchmarking Suite for Tool-Using Agents

2026-05-10 · Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu arxiv

Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims. We present an executable benchmarking suite that makes these objects explicit under a shared evidence-admission contract. The suite connects WebArena Verified, a SWE-Gym slice with SWE-bench-compatible verification, and MiniWoB++ through common workload adapters, task manifests, event schemas, replay/freeze policy, declared drivers, and reporting pipelines. In the canonical release, the gate separates paper-facing evidence from preflight, fixture, smoke, and diagnostic rows while preserving non-admitted artifacts for audit and onboarding. The admitted evidence records latency, invalid-action behavior, patch-generation cost, verifier metadata, replay bindings, and provenance under one auditable contract. The gate is decision-relevant rather than merely clerical: in a separate WebArena Verified controller study, clean-baseline and medium live-stressed evaluation select different fixed controller variants under the same workload and admission contract. The release is scoped as a benchmarking suite and admitted evidence, not a new agent policy, model leaderboard, backend comparison, or autonomous SWE-bench solver.

📄 PDF Abstract BibTeX arXiv:2605.11030

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

2025-10-24 · Jonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket 외 arxiv

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many su…

SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization

2026-06-04 · Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Shusen Liu 외 arxiv

Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. While general-purpose coding agents show strong capabilities, they of…

WirelessAgent++: Automated Agentic Workflow Design and Benchmarking for Wireless Networks

2026-02-28 · Jingwen Tong, Zijian Li, Fang Liu, Wei Guo 외 arxiv

The integration of large language models (LLMs) into wireless networks has sparked growing interest in building autonomous AI agents for wireless tasks. However, existing approaches rely heavily on manually crafted promp…

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

2026-06-24 · Yang Tian, Zhengpeng Shi, Yu Zhou, Bo Zhao arxiv

Large language models are increasingly deployed as agents that solve tasks by interacting with external tool environments. Although recent tool-use benchmarks increasingly cover complex task settings, they still largely …

MedCTA: A Benchmark for Clinical Tool Agents

2026-06-10 · Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker, Bernard Ghanem arxiv

To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Existing benchmarks largely evaluate isolated…

Question Answering