paper-with-me

홈 › Papers

REVEAL -- Reasoning and Evaluation of Visual Evidence through Aligned Language

2025-08-18 · Ipsita Praharaj, Yukta Butala, Badrikanath Praharaj, Yash Butala arxiv

The rapid advancement of generative models has intensified the challenge of detecting and interpreting visual forgeries, necessitating robust frameworks for image forgery detection while providing reasoning as well as localization. While existing works approach this problem using supervised training for specific manipulation or anomaly detection in the embedding space, generalization across domains remains a challenge. We frame this problem of forgery detection as a prompt-driven visual reasoning task, leveraging the semantic alignment capabilities of large vision-language models. We propose a framework, REVEAL (Reasoning and Evaluation of Visual Evidence through Aligned Language), that incorporates generalized guidelines. We propose two tangential approaches - (1) Holistic Scene-level Evaluation that relies on the physics, semantics, perspective, and realism of the image as a whole and (2) Region-wise anomaly detection that splits the image into multiple regions and analyzes each of them. We conduct experiments over datasets from different domains (Photoshop, DeepFake and AIGC editing). We compare the Vision Language Models against competitive baselines and analyze the reasoning provided by them.

📄 PDF Abstract BibTeX arXiv:2508.12543

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly DetectionVisual Reasoning

Similar Papers 제목 키워드 기반

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

2025-10-07 · Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman arxiv

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based ta…

Reinforcement Learning

CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning

2026-05-13 · Tengda Guo, Jie Leng, Hanlei Li, Yaoyuan Liang 외 arxiv

Vision-Language Models (VLMs) have achieved strong performance on general multimodal reasoning, yet remain challenged in integrating nonlocal visual information to support semantically underdetermined visual reasoning. W…

Multimodal ReasoningVisual Reasoning

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

2025-05-30 · Ai Jian, Weijie Qiu, Xiaokun Wang, Peiyu Wang 외

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …

DiagnosticImage Comprehensionvalid

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

2025-05-30 · Zhuobai Dong, Junchao Yi, Ziyuan Zheng, Haochen Han 외

Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GP…

Spatial ReasoningVisual Reasoning

VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

2026-05-20 · Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu, Lihong Chen 외 arxiv

Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) eval…

Visual Question AnsweringRelational ReasoningMultimodal Reasoning