paper-with-me

Papers

OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards

2026-03-19 · Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang, Jingjing Xie, Zhaoyang Liu, Zhoumianze Liu, Kaiming Jin, Jianze Liang, Zonglin Li, Feng Wu, Bowen Zhou, Zun Wang, Zichen Ding arxiv

Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet training is highly sensitive to the quality of the reward function. Existing reward approaches struggle to achieve both scalability and performance. To address this, we propose OS-Themis, a scalable and accurate multi-agent critic framework. Unlike a single judge, OS-Themis decomposes trajectories into verifiable milestones to isolate critical evidence for decision making and employs a review mechanism to strictly audit the evidence chain before making the final verdict. To facilitate evaluation, we further introduce OmniGUIRewardBench (OGRBench), a holistic cross-platform benchmark for GUI outcome rewards, where all evaluated models achieve their best performance under OS-Themis. Extensive experiments on AndroidWorld show that OS-Themis yields a 10.3% improvement when used to support online RL training, and a 6.9% gain when used for trajectory validation and filtering in the self-training loop, highlighting its potential to drive agent evolution.

📄 PDF Abstract BibTeX arXiv:2603.19191

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningDecision Making

Similar Papers 제목 키워드 기반

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

2026-06-23 · Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos arxiv

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii)…

Reinforcement Learning

Fairness Testing: Testing Software for Discrimination

2017-09-11 · Sainyam Galhotra, Yuriy Brun, Alexandra Meliou

This paper defines software fairness and discrimination and develops a testing-based method for measuring if and how much software discriminates, focusing on causality in discriminatory behavior. Evidence of software dis…

Fairnessvalid

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

2025-07-12 · Taolin Zhang, Maosong Cao, Alexander Lam, Songyang Zhang 외

Recently, the role of LLM-as-judge in evaluating large language models has gained prominence. However, current judge models suffer from narrow specialization and limited robustness, undermining their capacity for compreh…

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

2026-03-26 · Tzu-Yen Ma, Bo Zhang, Zichen Tang, Junpeng Ding 외 arxiv

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmark…

THEMIS: Towards Practical Intellectual Property Protection for Post-Deployment On-Device Deep Learning Models

2025-03-31 · Yujin Huang, Zhi Zhang, Qingchuan Zhao, Xingliang Yuan 외

On-device deep learning (DL) has rapidly gained adoption in mobile apps, offering the benefits of offline model inference and user privacy preservation over cloud-based approaches. However, it inevitably stores models on…

GPU