paper-with-me

홈 › Papers

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

2026-03-26 · Tzu-Yen Ma, Bo Zhang, Zichen Tang, Junpeng Ding, Haolin Tian, Yuanze Li, Zhuodi Hao, Zixin Ding, Zirui Wang, Xinyu Yu, Shiyao Peng, Yizhuo Zhao, Ruomeng Jiang, Yiling Huang, Peizhi Zhao, Jiayuan Chen, Weisheng Tan, Haocheng Gao, Yang Liu, Jiacheng Liu, Zhongjun Yang, Jiayu Huang, Haihong E arxiv

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces three major advances. (1) Real-World Scenarios and Complexity: Our benchmark comprises over 4,000 questions spanning seven scenarios, derived from authentic retracted-paper cases and carefully curated multimodal synthetic data. With 60.47% complex-texture images, THEMIS bridges the critical gap between existing benchmarks and the complexity of real-world academic fraud. (2) Fraud-Type Diversity and Granularity: THEMIS systematically covers five challenging fraud types and introduces 16 fine-grained manipulation operations. On average, each sample undergoes multiple stacked manipulation operations, with the diversity and difficulty of these manipulations demanding a high level of visual fraud reasoning from the models. (3) Multi-Dimensional Capability Evaluation: We establish a mapping from fraud types to five core visual fraud reasoning capabilities, thereby enabling an evaluation that reveals the distinct strengths and specific weaknesses of different models across these core capabilities. Experiments on 16 leading MLLMs show that even the best-performing model, GPT-5, achieves an overall performance of only 56.15%, demonstrating that our benchmark presents a stringent test. We expect THEMIS to advance the development of MLLMs for complex, real-world fraud reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2603.25089

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OS-Themis: A Scalable Critic Framework for Generalist GUI Rewards

2026-03-19 · Zehao Li, Zhenyu Wu, Yibo Zhao, Bowen Yang 외 arxiv

Reinforcement Learning (RL) has the potential to improve the robustness of GUI agents in stochastic environments, yet training is highly sensitive to the quality of the reward function. Existing reward approaches struggl…

Reinforcement LearningDecision Making

Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability

2024-06-26 · Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin 외

The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area. With the recent emergence of powerful large language models (LLMs), some studies have turned to LLM-based automat…

Language ModelingLanguage Modellingnlg evaluationText Generation

Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback

2026-06-23 · Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos arxiv

Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) transparency through explainability and (ii)…

Reinforcement Learning

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

2025-02-05 · Renjun Hu, Yi Cheng, Libin Meng, Jiaxin Xia 외

The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge that delivers sophisticated context-aware e…

Instruction FollowingKnowledge Distillation

Fairness Testing: Testing Software for Discrimination

2017-09-11 · Sainyam Galhotra, Yuriy Brun, Alexandra Meliou

This paper defines software fairness and discrimination and develops a testing-based method for measuring if and how much software discriminates, focusing on causality in discriminatory behavior. Evidence of software dis…

Fairnessvalid