paper-with-me

Papers

CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

2026-05-01 · Hang Wang, Chao Shen, Chenhao Lin, Minghui Yang, Lei Zhang, Cong Wang arxiv

The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video (AIGV) detection methods primarily focus on uni-modal or spatiotemporal artifacts, but they overlook the rich cues within the visual-textual cross-modal space, especially the temporal stability of semantic alignment. In this work, we identify a distinctive fingerprint in AIGVs, termed cross-modal temporal artifact (CMTA). Unlike real videos that exhibit natural temporal fluctuations in cross-modal alignment due to semantic variations, AIGVs display unnaturally stable semantic trajectories governed by given input prompts. To bridge this gap, we propose the CMTA framework, a cross-modal detection approach that captures these unique temporal artifacts through joint cross-modal embedding and multi-grained temporal modeling. Specifically, CMTA leverages BLIP to generate frame-level image captions and utilizes CLIP to extract corresponding visual-textual representations. A coarse-grained temporal modeling branch is then designed to characterize temporal fluctuations in cross-modal alignment with a GRU. In parallel, a fine-grained branch is constructed to capture intricate inter-frame variations from integrated visual-textual features with a Transformer encoder. Extensive experiments on 40 subsets across four large-scale datasets, including GenVideo, EvalCrafter, VideoPhy, and VidProM, validate that our approach sets a new state-of-the-art while exhibiting superior cross-generator generalization. Code and models of CMTA will be released at https://github.com/hwang-cs-ime/CMTA

📄 PDF Abstract BibTeX arXiv:2605.00630

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring

2024-08-27 · Taewoo Kim, Hoonhee Cho, Kuk-Jin Yoon

Video deblurring aims to enhance the quality of restored results in motion-blurred videos by effectively gathering information from adjacent video frames to compensate for the insufficient data in a single blurred frame.…

DeblurringVideo Deblurring

Looking for COVID-19 misinformation in multilingual social media texts

2021-05-03 · Raj Ratn Pranesh, Mehrdad Farokhnejad, Ambesh Shekhar, Genoveva Vargas-Solar

This paper presents the Multilingual COVID-19 Analysis Method (CMTA) for detecting and observing the spread of misinformation about this disease within texts. CMTA proposes a data science (DS) pipeline that applies machi…

Misinformation

Contrastive Modules with Temporal Attention for Multi-Task Reinforcement Learning

2023-11-02 · NeurIPS 2023 11

In the field of multi-task reinforcement learning, the modular principle, which involves specializing functionalities into different modules and combining them appropriately, has been widely adopted as a promising approa…

Contrastive LearningMulti-Task Learningreinforcement-learningReinforcement Learning

CMTA: COVID-19 Misinformation Multilingual Analysis on Twitter

2021-08-01 · ACL 2021 5 · Raj Pranesh, Mehrdad Farokhenajd, Ambesh Shekhar, Genoveva Vargas-Solar

The internet has actually come to be an essential resource of health knowledge for individuals around the world in the present situation of the coronavirus condition pandemic(COVID-19). During pandemic situations, myths,…

MisinformationRumour DetectionTransfer Learning

NP-TCMtarget: a network pharmacology platform for exploring mechanisms of action of Traditional Chinese medicine

2024-08-17 · Aoyi Wang, Yingdong Wang, Haoyang Peng, Haoran Zhang 외

The biological targets of traditional Chinese medicine (TCM) are the core effectors mediating the interaction between TCM and the human body. Identification of TCM targets is essential to elucidate the chemical basis and…

Binary Classification