paper-with-me

홈 › Papers

"Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents

2025-11-25 · Marta Sumyk, Oleksandr Kosovan arxiv

Computer Use Agents (CUAs) are designed to autonomously operate digital interfaces, yet they often fail to reliably determine whether a given task has been completed. We present an autonomous evaluation and feedback framework that uses vision-language models to assess task completion directly from screenshots and task descriptions. Our dataset covers 42 built-in macOS applications and 1,260 human-labeled tasks across a wide range of scenarios. Our framework achieves up to 73 percent accuracy in task success detection and yields an average relative improvement of 27 percent in overall task success when evaluator feedback is applied. These results show that vision-based evaluation can serve as an effective feedback mechanism that improves the reliability and self-correction of autonomous computer-use agents.

📄 PDF Abstract BibTeX arXiv:2511.20067

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PanDepth: Joint Panoptic Segmentation and Depth Completion

2022-12-29 · Juan Lagos, Esa Rahtu

Understanding 3D environments semantically is pivotal in autonomous driving applications where multiple computer vision tasks are involved. Multi-task models provide different types of outputs for a given scene, yielding…

Autonomous DrivingDepth CompletionInstance SegmentationPanoptic Segmentation+2

Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation

2026-06-23 · Marta Sumyk, Oleksandr Kosovan arxiv

Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environme…

Reinforcement Learning

Intelligent Mode-switching Framework for Teleoperation

2024-02-08 · Burak Kizilkaya, Changyang She, Guodong Zhao, Muhammad Ali Imran

Teleoperation can be very difficult due to limited perception, high communication latency, and limited degrees of freedom (DoFs) at the operator side. Autonomous teleoperation is proposed to overcome this difficulty by p…

Decision MakingDeep Reinforcement LearningIntent Detection

SemSegDepth: A Combined Model for Semantic Segmentation and Depth Completion

2022-09-01 · Juan Pablo Lagos, Esa Rahtu

Holistic scene understanding is pivotal for the performance of autonomous machines. In this paper we propose a new end-to-end model for performing semantic segmentation and depth completion jointly. The vast majority of …

Depth CompletionScene UnderstandingSegmentationSemantic Segmentation

Are All Point Clouds Suitable for Completion? Weakly Supervised Quality Evaluation Network for Point Cloud Completion

2023-03-03 · Jieqi Shi, Peiliang Li, Xiaozhi Chen, Shaojie Shen

In the practical application of point cloud completion tasks, real data quality is usually much worse than the CAD datasets used for training. A small amount of noisy data will usually significantly impact the overall sy…

AllAutonomous DrivingPoint Cloud Completion