paper-with-me

홈 › Papers

ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding

2024-12-17 · CVPR 2025 1 · Zhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong, Dan Guo, Meng Wang

We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for accurately manipulation detection and grounding. While existing DGM4 methods pay rare attention to the cross-modal alignment, hampering the accuracy of manipulation detecting to step further. To remedy this issue, this work targets to advance the semantic alignment learning to promote this task. Particularly, we utilize the off-the-shelf Multimodal Large-Language Models (MLLMs) and Large Language Models (LLMs) to construct paired image-text pairs, especially for the manipulated instances. Subsequently, a cross-modal alignment learning is performed to enhance the semantic alignment. Besides the explicit auxiliary clues, we further design a Manipulation-Guided Cross Attention (MGCA) to provide implicit guidance for augmenting the manipulation perceiving. With the grounding truth available during training, MGCA encourages the model to concentrate more on manipulated components while downplaying normal ones, enhancing the model's ability to capture manipulations. Extensive experiments are conducted on the DGM4 dataset, the results demonstrate that our model can surpass the comparison method with a clear margin.

📄 PDF Abstract BibTeX arXiv:2412.12718

Code (1)

CriliasMiller/ASAP 공식 구현 pytorch

Tasks

cross-modal alignment

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

ASAP: Automatic Semantic Alignment for Phrases

2014-08-01 · SEMEVAL 2014 8 · Ana Alves, Adriana Ferrugento, Mariana Louren{\c{c}}o, Filipe Rodrigues
Natural Language Inference

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

2026-05-30 · Rongsheng Wang, Fenghe Tang, Zihang Jiang, Yingtai Li 외 arxiv

Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak, heterogeneous supervision provided by radiology reports. In this pa…

Visual Question AnsweringRepresentation LearningCross-Modal Retrieval

Building Scalable Video Understanding Benchmarks through Sports

2023-01-17 · Aniket Agarwal, Alex Zhang, Karthik Narasimhan, Igor Gilitschenski 외

Existing benchmarks for evaluating long video understanding falls short on two critical aspects, either lacking in scale or quality of annotations. These limitations arise from the difficulty in collecting dense annotati…

Video Understanding

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

2025-06-25 · Yanzhe Chen, Huasong Zhong, Yan Li, Zhenheng Yang

Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existin…

16k

ASAP-II: From the Alignment of Phrases to Textual Similarity

2015-06-01 · SEMEVAL 2015 6 · Ana Alves, David Sim{\~o}es, Hugo Gon{\c{c}}alo Oliveira, Adriana Ferrugento
Natural Language InferenceSemantic Textual Similarity