VisCritic: Visual State Comparison as Process Reward for GUI Agents
GUI agents powered by vision-language models show strong potential for automating digital tasks, yet frequently fail in long-horizon scenarios due to the absence of step-level verification. Existing process reward models verify actions through textual reasoning alone, missing the visual nature of GUI state changes. We introduce VisCritic, a visual process reward framework that verifies agent actions by directly comparing pre-action and post-action screenshots in visual feature space. VisCritic employs a Siamese vision transformer to extract change-aware representations, coupled with an Action-Aware Critic Head that jointly evaluates action success, task progress, and error type. A critic-training data construction pipeline generates weakly supervised samples from existing trajectories without additional human labels for critic training. Experiments and offline analyses across five benchmarks demonstrate that VisCritic serves as a plug-and-play enhancement for diverse GUI agents, generally improving benchmark metrics while providing visual diagnostic cues.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Time-Efficient Reward Learning via Visually Assisted Cluster Ranking
One of the most successful paradigms for reward learning uses human feedback in the form of comparisons. Although these methods hold promise, human comparison labeling is expensive and time consuming, constituting a majo…
Dimensionality ReductionMuJoCoCompare and Select: Video Summarization with Multi-Agent Reinforcement Learning
Video summarization aims at generating concise video summaries from the lengthy videos, to achieve better user watching experience. Due to the subjectivity, purely supervised methods for video summarization may bring the…
Decision MakingMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning+3PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
In this paper, we make the first attempt to align diffusion models for image inpainting with human aesthetic standards via a reinforcement learning framework, significantly improving the quality and visual appeal of inpa…
3D ReconstructionImage Inpaintingreinforcement-learningReinforcement LearningSocial Comparison without Explicit Inference of Others' Reward Values: A Constructive Approach Using a Probabilistic Generative Model
Social comparison$\unicode{x2014}$the process of evaluating one's rewards relative to others$\unicode{x2014}$is an essential feature of social emotions such as envy and plays a fundamental role in primate social cognitio…
Learning rewards for robotic ultrasound scanning using probabilistic temporal ranking
Informative path-planning is a well established approach to visual-servoing and active viewpoint selection in robotics, but typically assumes that a suitable cost function or goal state is known. This work considers the …
Reinforcement Learning