paper-with-me

홈 › Papers

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

2025-06-06 · CVPR 2025 1 · Yiheng Li, Yang Yang, Zichang Tan, Huan Liu, Weihua Chen, Xu Zhou, Zhen Lei

To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation DGM4 has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local content, usually resulting in an inadequate perception of detailed forgery and unreliable results. In this paper, we propose a novel approach named Contextual-Semantic Consistency Learning (CSCL) to enhance the fine-grained perception ability of forgery for DGM4. Two branches for image and text modalities are established, each of which contains two cascaded decoders, i.e., Contextual Consistency Decoder (CCD) and Semantic Consistency Decoder (SCD), to capture within-modality contextual consistency and across-modality semantic consistency, respectively. Both CCD and SCD adhere to the same criteria for capturing fine-grained forgery details. To be specific, each module first constructs consistency features by leveraging additional supervision from the heterogeneous information of each token pair. Then, the forgery-aware reasoning or aggregating is adopted to deeply seek forgery cues based on the consistency features. Extensive experiments on DGM4 datasets prove that CSCL achieves new state-of-the-art performance, especially for the results of grounding manipulated content. Codes and weights are avaliable at https://github.com/liyih/CSCL.

📄 PDF Abstract BibTeX arXiv:2506.05890

Code (1)

liyih/cscl 공식 구현 pytorch

Tasks

Decoder

Similar Papers 제목 키워드 기반

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

2025-09-18 · Zaiquan Yang, Yuhao Liu, Gerhard Hancke, Rynson W. H. Lau arxiv

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-sh…

Spatio-Temporal Video Grounding

Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding

2025-08-04 · Xinquan Yu, Wei Lu, Xiangyang Luo arxiv

The task of Detecting and Grounding Multi-Modal Media Manipulation (DGM$^4$) is a branch of misinformation detection. Unlike traditional binary classification, it includes complex subtasks such as forgery content localiz…

Binary Classification

On the Consistency of Video Large Language Models in Temporal Comprehension

2024-11-20 · CVPR 2025 1 · Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang, Angela Yao

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on …

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

2025-10-04 · Mingyu Liu, Zheng Huang, Xiaoyi Lin, Muzhi Zhu 외 arxiv

Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-Language-Action methods typically entang…

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding

2025-05-17 · Jingkun Yue, Siqi Zhang, Zinan Jia, Huihuan Xu 외

Visual grounding is essential for precise perception and reasoning in multimodal large language models (MLLMs), especially in medical imaging domains. While existing medical visual grounding benchmarks primarily focus on…

Visual GroundingVisual Question Answering (VQA)