paper-with-me

홈 › Papers

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark

2025-07-22 · Xu Yang, Qi Zhang, Shuming Jiang, Yaowen Xu, Zhaofan Zou, Hao Sun, Xuelong Li arxiv

With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing detection approaches predominantly focus on binary classification while lacking detailed and interpretable explanations of forgeries, which limits their applicability in safety-critical scenarios. Moreover, current methods often treat each modality separately, without a unified benchmark for cross-modal forgery detection and interpretation. To address these challenges, we introduce METER, a unified, multi-modal benchmark for interpretable forgery detection spanning images, videos, audio, and audio-visual content. Our dataset comprises four tracks, each requiring not only real-vs-fake classification but also evidence-chain-based explanations, including spatio-temporal localization, textual rationales, and forgery type tracing. Compared to prior benchmarks, METER offers broader modality coverage and richer interpretability metrics such as spatial/temporal IoU, multi-class tracing, and evidence consistency. We further propose a human-aligned, three-stage Chain-of-Thought (CoT) training strategy combining SFT, DPO, and a novel GRPO stage that integrates a human-aligned evaluator with CoT reasoning. We hope METER will serve as a standardized foundation for advancing generalizable and interpretable forgery detection in the era of generative media.

📄 PDF Abstract BibTeX arXiv:2507.16206

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Similar Papers 제목 키워드 기반

Real-Time Visual Attribution Streaming in Thinking Model

2026-04-17 · Seil Kang, Woojung Han, Junhyeok Kim, Jinyeong Kim 외 arxiv

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning trace…

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

2025-09-25 · Tianrun Xu, Haoda Jing, Ye Li, Yuquan Wei 외 arxiv

Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing methods still fail to ensure evidence-answer consistency, where correct…

Reinforcement LearningMultimodal Reasoning

REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection

2025-11-28 · Huangsen Cao, Qin Mei, Zhiheng Li, Yuxi Li 외 arxiv

The rapid progress of visual generative models has made AI-generated images increasingly difficult to distinguish from authentic ones, posing growing risks to social trust and information integrity. This motivates detect…

Reinforcement LearningDomain Generalization

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

2026-02-10 · Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang 외 arxiv

Verifying the truthfulness of claims usually requires joint multi-modal reasoning over both textual and visual evidence, such as analyzing both textual caption and chart image for claim verification. In addition, to make…

Explanation Generation

Multimodal Explanations: Justifying Decisions and Pointing to the Evidence

2018-02-15 · CVPR 2018 6 · Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach 외

Deep models that are both effective and explainable are desirable in many settings; prior explainable models have been unimodal, offering either image-based visualization of attention weights or text-based generation of …

Activity RecognitionExplainable ModelsQuestion AnsweringVisual Question Answering+1