"Are We Done Yet?": A Vision-Based Judge for Autonomous Task Completion of Computer Use Agents
Computer Use Agents (CUAs) are designed to autonomously operate digital interfaces, yet they often fail to reliably determine whether a given task has been completed. We present an autonomous evaluation and feedback framework that uses vision-language models to assess task completion directly from screenshots and task descriptions. Our dataset covers 42 built-in macOS applications and 1,260 human-labeled tasks across a wide range of scenarios. Our framework achieves up to 73 percent accuracy in task success detection and yields an average relative improvement of 27 percent in overall task success when evaluator feedback is applied. These results show that vision-based evaluation can serve as an effective feedback mechanism that improves the reliability and self-correction of autonomous computer-use agents.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
PanDepth: Joint Panoptic Segmentation and Depth Completion
Understanding 3D environments semantically is pivotal in autonomous driving applications where multiple computer vision tasks are involved. Multi-task models provide different types of outputs for a given scene, yielding…
Autonomous DrivingDepth CompletionInstance SegmentationPanoptic Segmentation+2Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation
Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environme…
Reinforcement LearningIntelligent Mode-switching Framework for Teleoperation
Teleoperation can be very difficult due to limited perception, high communication latency, and limited degrees of freedom (DoFs) at the operator side. Autonomous teleoperation is proposed to overcome this difficulty by p…
Decision MakingDeep Reinforcement LearningIntent DetectionSemSegDepth: A Combined Model for Semantic Segmentation and Depth Completion
Holistic scene understanding is pivotal for the performance of autonomous machines. In this paper we propose a new end-to-end model for performing semantic segmentation and depth completion jointly. The vast majority of …
Depth CompletionScene UnderstandingSegmentationSemantic SegmentationAre All Point Clouds Suitable for Completion? Weakly Supervised Quality Evaluation Network for Point Cloud Completion
In the practical application of point cloud completion tasks, real data quality is usually much worse than the CAD datasets used for training. A small amount of noisy data will usually significantly impact the overall sy…
AllAutonomous DrivingPoint Cloud Completion