paper-with-me

홈 › Papers

Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization

2025-08-25 · Keyang Zhang, Chenqi Kong, Hui Liu, Bo Ding, Xinghao Jiang, Haoliang Li arxiv

The increasing sophistication of image manipulation techniques demands robust forensic solutions that can both reliably detect alterations and precisely localize tampered regions. Recent Multimodal Large Language Models (MLLMs) show promise by leveraging world knowledge and semantic understanding for context-aware detection, yet they struggle with perceiving subtle, low-level forensic artifacts crucial for accurate manipulation localization. This paper presents a novel Propose-Rectify framework that effectively bridges semantic reasoning with forensic-specific analysis. In the proposal stage, our approach utilizes a forensic-adapted LLaVA model to generate initial manipulation analysis and preliminary localization of suspicious regions based on semantic understanding and contextual reasoning. In the rectification stage, we introduce a Forensics Rectification Module that systematically validates and refines these initial proposals through multi-scale forensic feature analysis, integrating technical evidence from several specialized filters. Additionally, we present an Enhanced Segmentation Module that incorporates critical forensic cues into SAM's encoded image embeddings, thereby overcoming inherent semantic biases to achieve precise delineation of manipulated regions. By synergistically combining advanced multimodal reasoning with established forensic methodologies, our framework ensures that initial semantic proposals are systematically validated and enhanced through concrete technical evidence, resulting in comprehensive detection accuracy and localization precision. Extensive experimental validation demonstrates state-of-the-art performance across diverse datasets with exceptional robustness and generalization capabilities.

📄 PDF Abstract BibTeX arXiv:2508.17976

Code (0)

등록된 구현이 없습니다.

Tasks

Image Manipulation LocalizationMultimodal Reasoning

Similar Papers 제목 키워드 기반

ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models

2025-12-01 · Xusen Hei, Jiali Chen, Jinyu Yang, Mengchen Zhao 외 arxiv

As multimodal large language models (MLLMs) frequently exhibit errors in complex video reasoning scenarios, correcting these errors is critical for uncovering their weaknesses and improving performance. However, existing…

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

2026-07-03 · Geng Li, Yuxin Peng arxiv

While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, particularly for tiny objects in cluttered sce…

PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs

2026-08-19 · Yaqi Li, Jielun Peng, Yabin Wang, Jincheng Liu 외 arxiv

Existing explainable deepfake forensic methods typically rely on task-adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent-style tool use, we instead introduce a Perception-as-Tool …

Explanation Generation

ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization

2026-06-23 · Lei Xu, Haowei Wang, Shen Chen, Taiping Yao 외 arxiv

Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptimal localization. The reliance on stitched…

Image Manipulation Localization

BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM

2025-07-19 · Haiquan Wen, Tianxiao Li, Zhenglin Huang, Yiwei He 외 arxiv

The rapid advancement of generative AI has substantially improved image and video synthesis, amplifying the risk of multimodal visual misinformation. Recent MLLMs have shown promise for transparent AI-generated content d…

Visual Reasoning