paper-with-me

홈 › Papers

MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents

2026-06-02 · Jia Yu, Zilong Wang, Xinyang Jiang, Dongsheng Li, Shuo Wang arxiv

Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated. Existing benchmarks focus on general web or desktop tasks and underrepresent medical software, which requires domain knowledge, exhibits markedly different UI design from mainstream applications, lacks public testing environments, and demands safety validation beyond task completion. We introduce MedCUA-Bench, an interactive benchmark for clinical computer-use agents. It covers 18 clinical scenarios across 10 medical domains, reconstructed from real product manuals and open-source medical systems to capture authentic clinical interfaces while avoiding licensing and privacy constraints. Each task ships with paired intent- and step-level goals to disentangle clinical reasoning from UI execution, and is evaluated by a deterministic checker over task completion and five clinical safety dimensions. Across 23 agents, the best closed-source model reaches 54.2% strict success, while all models remain below 9% on the real OpenEMR. Open-source agents average only 2.5%, with the best reaching 16.2%. MedCUA-Bench exposes the gap between current agents and reliable clinical software use, providing a reproducible testbed for future research.

📄 PDF Abstract BibTeX arXiv:2606.03203

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents

2026-05-22 · JunJia Guo, Yuhang Yao, Jiawei, Zhou 외 arxiv

We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike prior code generation benchmarks that focus on algorithmic tasks, VIS…

Code Generation

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

2026-07-07 · Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen 외 arxiv

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expres…

OmniParser for Pure Vision Based GUI Agent

2024-08-01

The recent success of large vision language models shows great potential in driving the agent system operating on user interfaces. However, we argue that the power multimodal models like GPT-4V as a general agent on mult…

Natural Language Visual Grounding

Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval

2025-02-17 · Ze Liu, Zhengyang Liang, Junjie Zhou, Zheng Liu 외

With the popularity of multimodal techniques, it receives growing interests to acquire useful information in visual forms. In this work, we formally define an emerging IR paradigm called \textit{Visualized Information Re…

Information RetrievalRetrieval

ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions

2026-08-27 · Rui Xie, Lu Chen arxiv

Powerful code agents can execute scripts, call tools, and manage files, yet many important applications remain accessible primarily through graphical user interfaces. We argue that screenshot-and-click is an inefficient …