paper-with-me

홈 › Papers

A Large-scale Interpretable Multi-modality Benchmark for Facial Image Forgery Localization

2024-12-27 · Jingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu, Li Zhu, Zhedong Zheng

Image forgery localization, which centers on identifying tampered pixels within an image, has seen significant advancements. Traditional approaches often model this challenge as a variant of image segmentation, treating the binary segmentation of forged areas as the end product. We argue that the basic binary forgery mask is inadequate for explaining model predictions. It doesn't clarify why the model pinpoints certain areas and treats all forged pixels the same, making it hard to spot the most fake-looking parts. In this study, we mitigate the aforementioned limitations by generating salient region-focused interpretation for the forgery images. To support this, we craft a Multi-Modal Tramper Tracing (MMTT) dataset, comprising facial images manipulated using deepfake techniques and paired with manual, interpretable textual annotations. To harvest high-quality annotation, annotators are instructed to meticulously observe the manipulated images and articulate the typical characteristics of the forgery regions. Subsequently, we collect a dataset of 128,303 image-text pairs. Leveraging the MMTT dataset, we develop ForgeryTalker, an architecture designed for concurrent forgery localization and interpretation. ForgeryTalker first trains a forgery prompter network to identify the pivotal clues within the explanatory text. Subsequently, the region prompter is incorporated into multimodal large language model for finetuning to achieve the dual goals of localization and interpretation. Extensive experiments conducted on the MMTT dataset verify the superior performance of our proposed model. The dataset, code as well as pretrained checkpoints will be made publicly available to facilitate further research and ensure the reproducibility of our results.

📄 PDF Abstract BibTeX arXiv:2412.19685

Code (0)

등록된 구현이 없습니다.

Tasks

Face SwappingImage SegmentationLarge Language ModelMultimodal Large Language ModelSemantic Segmentation

Similar Papers 제목 키워드 기반

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

2026-05-25 · Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman M. Shaker 외 arxiv

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusio…

Visual Reasoning

Judge Model for Large-scale Multimodality Benchmarks

2026-01-03 · Min-Han Shih, Yu-Hsin Wu, Yu-Wei Chen arxiv

We propose a dedicated multimodal Judge Model designed to provide reliable, explainable evaluation across a diverse suite of tasks. Our benchmark spans text, audio, image, and video modalities, drawing from carefully sam…

Interpretable multimodal sentiment analysis based on textual modality descriptions by using large-scale language models

2023-05-07 · Sixia Li, Shogo Okada

Multimodal sentiment analysis is an important area for understanding the user's internal states. Deep learning methods were effective, but the problem of poor interpretability has gradually gained attention. Previous wor…

Multimodal Sentiment AnalysisSentiment Analysis

Conditional Evidence Reconstruction and Decomposition for Interpretable Multimodal Diagnosis

2026-04-18 · Shaowen Wan, Yanjun Lv, Lu Zhang, Dajiang Zhu 외 arxiv

Neurobiological and neurodegenerative diseases are inherently multifactorial, arising from coupled influences spanning genetic susceptibility, brain alterations, and environmental and behavioral factors. Multimodal model…

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

2026-03-19 · Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 외 arxiv

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this pro…

Temporal SequencesSpatial ReasoningVisual Reasoning