paper-with-me

Papers

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

2026-07-05 · Shen Shen, Quan Zhang, Dan Jiang, Ke Zhang arxiv

The rapid proliferation of artificial intelligence-generated content necessitates reliable multimodal forensics. Beyond video-level binary classification, precisely localizing sparsely distributed forged segments in long-form videos remains a critical challenge. This task is particularly difficult when manipulations are subtly embedded and cross-modal signals are weak and temporally diffuse. To address these challenges, we propose EVAS, an end-to-end multimodal framework for temporal forgery localization. At its core, a Multi-Stage Audio-Visual Synergy mechanism facilitates progressive cross-modal interaction to learn deep multimodal forensic representations and capture high-order semantic traces of sparse manipulations. Furthermore, we introduce a Boundary-Aware Refinement strategy to achieve steered boundary calibration. By incorporating invalid-frame masking, this strategy suppresses ambiguous regions and sharpens transition predictions. We adopt a decoupled training paradigm with auxiliary heads to disentangle representation learning from inference objectives, enhancing model generalization and stability. Additionally, a lightweight HourglassFFN is incorporated to reduce computational overhead. Extensive experiments demonstrate that EVAS achieves state-of-the-art average localization accuracy and average recall across three benchmark datasets, validating its effectiveness for fine-grained temporal forgery localization.

📄 PDF Abstract BibTeX arXiv:2607.04472

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningBinary Classification

Similar Papers 제목 키워드 기반

Do You Really Mean That? Content Driven Audio-Visual Deepfake Dataset and Multimodal Method for Temporal Forgery Localization

2022-04-13 · Zhixi Cai, Kalin Stefanov, Abhinav Dhall, Munawar Hayat

Due to its high societal impact, deepfake detection is getting active attention in the computer vision community. Most deepfake detection methods rely on identity, facial attributes, and adversarial perturbation-based sp…

BenchmarkingDeepFake DetectionTemporal Forgery Localization

A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization

2025-07-22 · Wenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo 외 arxiv

Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and challenging to scale for large datasets…

Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning

2025-08-04 · Wenbo Xu, Wei Lu, Xiangyang Luo arxiv

The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detection and localization, there is still a la…

Binary ClassificationDeepFake Detection

Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network

2025-05-03 · Junyan Wu, Wenbo Xu, Wei Lu, Xiangyang Luo 외

Audio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained a…

Contrastive LearningTemporal Forgery Localization

Glitch in the Matrix: A Large Scale Benchmark for Content Driven Audio-Visual Forgery Detection and Localization

2023-05-03 · Zhixi Cai, Shreya Ghosh, Abhinav Dhall, Tom Gedeon 외

Most deepfake detection methods focus on detecting spatial and/or spatio-temporal changes in facial attributes and are centered around the binary classification task of detecting whether a video is real or fake. This is …

Binary ClassificationDeepFake DetectionFace SwappingTemporal Forgery Localization