paper-with-me

Papers

PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies

2025-10-18 · Lukas Selch, Yufang Hou, M. Jehanzeb Mirza, Sivan Doveh, James Glass, Rogerio Feris, Wei Lin arxiv

Large Multimodal Models (LMMs) are increasingly applied to scientific research, yet it remains unclear whether they can reliably understand and reason over the multimodal complexity of papers. A central challenge lies in detecting and resolving inconsistencies across text, figures, tables, and equations, issues that are often subtle, domain-specific, and ultimately undermine clarity, reproducibility, and trust. Existing benchmarks overlook this issue, either isolating single modalities or relying on synthetic errors that fail to capture real-world complexity. We introduce PRISMM-Bench (Peer-Review-sourced Inconsistency Set for Multimodal Models), the first benchmark grounded in real reviewer-flagged inconsistencies in scientific papers. Through a multi-stage pipeline of review mining, LLM-assisted filtering and human verification, we curate 384 inconsistencies from 353 papers. Based on this set, we design three tasks, namely inconsistency identification, remedy and pair matching, which assess a model's capacity to detect, correct, and reason over inconsistencies across different modalities. Furthermore, to address the notorious problem of choice-only shortcuts in multiple-choice evaluation, where models exploit answer patterns without truly understanding the question, we further introduce structured JSON-based answer representations that minimize linguistic biases by reducing reliance on superficial stylistic cues. We benchmark 21 leading LMMs, including large open-weight models (GLM-4.5V 106B, InternVL3 78B) and proprietary models (Gemini 2.5 Pro, GPT-5 with high reasoning). Results reveal strikingly low performance (27.8-53.9\%), underscoring the challenge of multimodal scientific reasoning and motivating progress towards trustworthy scientific assistants.

📄 PDF Abstract BibTeX arXiv:2510.16505

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PeerRank: Autonomous LLM Evaluation Through Web-Grounded, Bias-Controlled Peer Review

2026-02-01 · Yanki Margalit, Erni Avram, Ran Taig, Oded Margalit 외 arxiv

Evaluating large language models typically relies on human-authored benchmarks, reference answers, and human or single-model judgments, approaches that scale poorly, become quickly outdated, and mismatch open-world deplo…

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

2026-07-12 · Nuo Chen, Qian Wang, Qingyun Zou, Bingsheng He arxiv

When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We opera…

CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?

2025-03-27 · Jiefu Ou, William Gantt Walden, Kate Sanders, Zhengping Jiang 외

A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically generate plausible (if generic) reviews, ensur…

BenchmarkingSpecificity

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

2026-04-05 · Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang 외 arxiv

LLM-based reviewing systems typically take only the manuscript as input, leaving literature and code-based claims hard to verify. We present FactReview, a system that extracts review-relevant claims, grounds them in rela…

Is Your Paper Being Reviewed by an LLM? A New Benchmark Dataset and Approach for Detecting AI Text in Peer Review

2025-02-26 · Sungduk Yu, Man Luo, Avinash Madusu, Vasudev Lal 외

Peer review is a critical process for ensuring the integrity of published scientific research. Confidence in this process is predicated on the assumption that experts in the relevant domain give careful consideration to …

BenchmarkingText Detection