paper-with-me

Papers

MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation

2025-08-19 · Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Ting Liu, Yuzhuo Fu arxiv

With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) have increasingly been adopted to assist in the generation of review comments; however, current LLM-based review tasks lack a unified evaluation benchmark to rigorously assess the models' ability to produce comprehensive, accurate, and human-aligned assessments, particularly in scenarios involving multimodal content such as figures and tables. To address this gap, we propose \textbf{MMReview}, a comprehensive benchmark that spans multiple disciplines and modalities. MMReview includes multimodal content and expert-written review comments for 240 papers across 17 research domains within four major academic disciplines: Artificial Intelligence, Natural Sciences, Engineering Sciences, and Social Sciences. We design a total of 13 tasks grouped into four core categories, aimed at evaluating the performance of LLMs and Multimodal LLMs (MLLMs) in step-wise review generation, outcome formulation, alignment with human preferences, and robustness to adversarial input manipulation. Extensive experiments conducted on 16 open-source models and 5 advanced closed-source models demonstrate the thoroughness of the benchmark. We envision MMReview as a critical step toward establishing a standardized foundation for the development of automated peer review systems.

📄 PDF Abstract BibTeX arXiv:2508.14146

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FMMD: A multimodal open peer review dataset based on F1000Research

2026-02-15 · Zhenzhen Zhuang, Yuqing Fu, Jing Zhu, Zhangping Zhou 외 arxiv

Automated scholarly paper review (ASPR) has entered the coexistence phase with traditional peer review, where artificial intelligence (AI) systems are increasingly incorporated into real-world manuscript evaluation. In p…

MOPRD: A multidisciplinary open peer review dataset

2022-12-09 · Jialiang Lin, Jiaxin Song, Zhangping Zhou, Yidong Chen 외

Open peer review is a growing trend in academic publications. Public access to peer review data can benefit both the academic and publishing communities. It also serves as a great support to studies on review comment gen…

Comment GenerationReview Generation

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

2026-06-10 · Xinyu Zhao, Rana Muhammad Shahroz Khan, Zhen Xu, Zhen Tan 외 arxiv

The integration of Large Language Models (LLMs) and Multimodal LLMs (MLLMs) into scientific peer-review workflows introduces novel and significant risks for adversarial manipulation, especially given the multimodal natur…

Proceedings of the 20th International Conference on Knowledge, Information and Creativity Support Systems (KICSS 2025)

2025-11-28 · Edited by Tessai Hayama, Takayuki Ito, Takahiro Uchiya, Motoki Miura 외 arxiv

This volume presents the proceedings of the 20th International Conference on Knowledge, Information and Creativity Support Systems (KICSS 2025), held in Nagaoka, Japan, on December 3-5, 2025. The conference, organized in…

MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

2024-07-06 · Zekun Li, Xianjun Yang, Kyuri Choi, Wanrong Zhu 외

The rapid development of Multimodal Large Language Models (MLLMs) is making AI-driven scientific assistants increasingly feasible, with interpreting scientific figures being a crucial task. However, existing datasets and…

ArticlesInstruction FollowingMultiple-choicevisual instruction following