paper-with-me

홈 › Papers

CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents

2026-03-11 · Marta Sumyk, Oleksandr Kosovan arxiv

Computer-Use Agents (CUAs) are emerging as a new paradigm in human-computer interaction, enabling autonomous execution of tasks in desktop environment by perceiving high-level natural-language instructions. As such agents become increasingly capable and are deployed across diverse desktop environments, evaluating their behavior in a scalable and reliable manner becomes a critical challenge. Existing evaluation pipelines rely on static benchmarks, rule-based success checks, or manual inspection, which are brittle, costly, and poorly aligned with real-world usage. In this work, we study Vision-Language Models (VLMs) as autonomous auditors for assessing CUA task completion directly from observable interactions and conduct a large-scale meta-evaluation of five VLMs that judge task success given a natural-language instruction and the final environment state. Our evaluation spans three widely used CUA benchmarks across macOS, Windows, and Linux environments and analyzes auditor behavior along three complementary dimensions: accuracy, calibration of confidence estimates, and inter-model agreement. We find that while state-of-the-art VLMs achieve strong accuracy and calibration, all auditors exhibit notable performance degradation in more complex or heterogeneous environments, and even high-performing models show significant disagreement in their judgments. These results expose fundamental limitations of current model-based auditing approaches and highlight the need to explicitly account for evaluator reliability, uncertainty, and variance when deploying autonomous CUAs in real-world settings.

📄 PDF Abstract BibTeX arXiv:2603.10577

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Black-Box Access is Insufficient for Rigorous AI Audits

2024-01-25 · Stephen Casper, Carson Ezell, Charlotte Siegmann, Noam Kolt 외

External audits of AI systems are increasingly recognized as a key mechanism for AI governance. The effectiveness of an audit, however, depends on the degree of access granted to auditors. Recent audits of state-of-the-a…

Law and the Emerging Political Economy of Algorithmic Audits

2024-04-03 · Petros Terzis, Michael Veale, Noëlle Gaumann

For almost a decade now, scholarship in and beyond the ACM FAccT community has been focusing on novel and innovative ways and methodologies to audit the functioning of algorithmic systems. Over the years, this research i…

On Learning and Enforcing Latent Assessment Models using Binary Feedback from Human Auditors Regarding Black-Box Classifiers

2022-02-16 · Mukund Telukunta, Venkata Sriram Siddhardh Nadendla

Algorithmic fairness literature presents numerous mathematical notions and metrics, and also points to a tradeoff between them while satisficing some or all of them simultaneously. Furthermore, the contextual nature of f…

FairnessPAC learning

GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions

2024-06-21 · Giulio Salierno, Rosamaria Bertè, Luca Attias, Carla Morrone 외

Recent advances in Natural Language Processing have demonstrated the effectiveness of pretrained language models like BERT for a variety of downstream tasks. We present GiusBERTo, the first BERT-based model specialized f…

De-identificationLanguage ModelingLanguage Modelling

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

2026-04-17 · Eric Gan, Aryan Bhatt, Buck Shlegeris, Julian Stastny 외 arxiv

As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results while evading detection. We introduce Auditing Sabotage Bench, a bench…