paper-with-me

홈 › Papers

M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency

2026-04-01 · Abolfazl Ansari, Delvin Ce Zhang, Zhuoyang Zou, Wenpeng Yin, Dongwon Lee arxiv

Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing benchmarks lack the scale, domain diversity, and visual complexity needed to evaluate this alignment realistically. To address this gap, we introduce M2-Verify, a large-scale multimodal dataset for checking scientific claim consistency. Sourced from PubMed and arXiv, M2-Verify provides over 469K instances across 16 domains, rigorously validated through expert audits. Extensive baseline experiments show that state-of-the-art models struggle to maintain robust consistency. While top models achieve up to 85.8\% Micro-F1 on low-complexity medical perturbations, performance drops to 61.6\% on high-complexity challenges like anatomical shifts. Furthermore, expert evaluations expose hallucinations when models generate scientific explanations for their alignment decisions. Finally, we demonstrate our dataset's utility and provide comprehensive usage guidelines.

📄 PDF Abstract BibTeX arXiv:2604.01306

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Turn-based Multi-Agent Reinforcement Learning Model Checking

2025-01-06 · Dennis Gross

In this paper, we propose a novel approach for verifying the compliance of turn-based multi-agent reinforcement learning (TMARL) agents with complex requirements in stochastic multiplayer games. Our method overcomes the …

modelMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking

2026-01-13 · Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach arxiv

The growing scale of online misinformation urgently demands Automated Fact-Checking (AFC). Existing benchmarks for evaluating AFC systems, however, are largely limited in terms of task scope, modalities, domain, language…

CrowdChecked: Detecting Previously Fact-Checked Claims in Social Media

2022-10-10 · Momchil Hardalov, Anton Chernyavskiy, Ivan Koychev, Dmitry Ilvovsky 외

While there has been substantial progress in developing systems to automate fact-checking, they still lack credibility in the eyes of the users. Thus, an interesting approach has emerged: to perform automatic fact-checki…

Fact Checking

Improving Evidence Retrieval for Automated Explainable Fact-Checking

2021-06-01 · NAACL 2021 4 · Chris Samarinas, Wynne Hsu, Mong Li Lee

Automated fact-checking on a large-scale is a challenging task that has not been studied systematically until recently. Large noisy document collections like the web or news articles make the task more difficult. We desc…

ArticlesFact CheckingRetrievalSentence

CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval

2026-01-23 · Akshith Reddy Putta, Jacob Devasier, Chengkai Li arxiv

Automated Fact-Checking has largely focused on verifying general knowledge against static corpora, overlooking high-stakes domains like law where truth is evolving and technically complex. We introduce CaseFacts, a bench…

Semantic SimilarityGeneral KnowledgeFact Verification