paper-with-me

Papers

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

2026-06-23 · Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos arxiv

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii) alignment via human feedback. While both show promising results, no publicly available framework currently combines them. To address this, we introduce Themis, an XAI-enabled testing and evaluation framework for Reinforcement Learning from Human Feedback. Themis supports over 200 widely used environments and is easily configurable for experiments in RL, transparency, and alignment. Our results show that Themis can train reward models that match or outperform the environment's true reward signal using human preferences. We also provide a cloud-based platform for collecting human feedback and managing experiments. It is user-friendly, auto-scalable, and supports large participant groups across multiple experiments without extra development overhead. Tests show Themis can support one thousand users in back-to-back experiments on a modest commercial machine.

📄 PDF Abstract BibTeX arXiv:2606.24622

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards

2026-03-19 · Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang 외 arxiv

Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet training is highly sensitive to the quality of the reward function. Existing reward approaches struggl…

Reinforcement LearningDecision Making

THEMIS: Towards Practical Intellectual Property Protection for Post-Deployment On-Device Deep Learning Models

2025-03-31 · Yujin Huang, Zhi Zhang, Qingchuan Zhao, Xingliang Yuan 외

On-device deep learning (DL) has rapidly gained adoption in mobile apps, offering the benefits of offline model inference and user privacy preservation over cloud-based approaches. However, it inevitably stores models on…

GPU

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

2025-02-05 · Renjun Hu, Yi Cheng, Libin Meng, Jiaxin Xia 외

The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge that delivers sophisticated context-aware e…

Instruction FollowingKnowledge Distillation

Fairness Testing: Testing Software for Discrimination

2017-09-11 · Sainyam Galhotra, Yuriy Brun, Alexandra Meliou

This paper defines software fairness and discrimination and develops a testing-based method for measuring if and how much software discriminates, focusing on causality in discriminatory behavior. Evidence of software dis…

Fairnessvalid

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

2026-03-26 · Tzu-Yen Ma, Bo Zhang, Zichen Tang, Junpeng Ding 외 arxiv

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmark…