paper-with-me

Papers

Reduced Spatial Dependency for More General Video-level Deepfake Detection

2025-03-05 · Beilin Chu, Xuan Xu, Yufei Zhang, Weike You, Linna Zhou

As one of the prominent AI-generated content, Deepfake has raised significant safety concerns. Although it has been demonstrated that temporal consistency cues offer better generalization capability, existing methods based on CNNs inevitably introduce spatial bias, which hinders the extraction of intrinsic temporal features. To address this issue, we propose a novel method called Spatial Dependency Reduction (SDR), which integrates common temporal consistency features from multiple spatially-perturbed clusters, to reduce the dependency of the model on spatial information. Specifically, we design multiple Spatial Perturbation Branch (SPB) to construct spatially-perturbed feature clusters. Subsequently, we utilize the theory of mutual information and propose a Task-Relevant Feature Integration (TRFI) module to capture temporal features residing in similar latent space from these clusters. Finally, the integrated feature is fed into a temporal transformer to capture long-range dependencies. Extensive benchmarks and ablation studies demonstrate the effectiveness and rationale of our approach.

📄 PDF Abstract BibTeX arXiv:2503.03270

Code (0)

등록된 구현이 없습니다.

Tasks

DeepFake DetectionFace Swapping

Similar Papers 제목 키워드 기반

Exemplar-based Video Colorization with Long-term Spatiotemporal Dependency

2023-03-27 · Siqi Chen, Xueming Li, Xianlin Zhang, Mingdao Wang 외

Exemplar-based video colorization is an essential technique for applications like old movie restoration. Although recent methods perform well in still scenes or scenes with regular movement, they always lack robustness i…

Colorization

Context Guided Transformer Entropy Modeling for Video Compression

2025-08-03 · Junlong Tong, Wei Zhang, Yaohui Jin, Xiaoyu Shen arxiv

Conditional entropy models effectively leverage spatio-temporal contexts to reduce video redundancy. However, incorporating temporal context often introduces additional model complexity and increases computational cost. …

Frequency-Aware Spatiotemporal Transformers for Video Inpainting Detection

2021-01-01 · ICCV 2021 10 · Bingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 외

In this paper, we propose a frequency-aware spatiotemporal transformers for deep In this paper, we propose a Frequency-Aware Spatiotemporal Transformer (FAST) for video inpainting detection, which aims to simultaneou…

DecoderVideo Inpainting

UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation Learning

2021-09-29 · ICLR 2022 4 · Kunchang Li, Yali Wang, Gao Peng, Guanglu Song 외

It is a challenging task to learn rich and multi-scale spatial-temporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between video frames. The recent advances in thi…

Action ClassificationAction RecognitionRepresentation Learning

Video-Based Human Pose Regression via Decoupled Space-Time Aggregation

2024-03-29 · CVPR 2024 1 · Jijie He, Wenwu Yang

By leveraging temporal dependency in video sequences, multi-frame human pose estimation algorithms have demonstrated remarkable results in complicated situations, such as occlusion, motion blur, and video defocus. These …

Pose Estimationregression