paper-with-me

홈 › Papers

Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection

2026-05-21 · Zhengcen Li, Chenyang Jiang, Liangxu Su, Tong Shao, Shiyang Zhou, Ming Tao, Jingyong Su arxiv

AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and visual modalities. A strongly generalizable detector should remain robust under distributional variations. However, we identify a consistent failure mode: SOTA AI-generated image detectors often collapse when applied to frames extracted from videos. Through systematic analysis, we show that this cross-modal gap arises from both entangled synthesis-agnostic video processing shifts, including color conversion, codec compression, resizing, and blur, and model-specific fingerprints introduced by modern video generators. Motivated by these findings, we propose VINA (Video as Natural Augmentation), a unified AIGC detection framework that jointly trains on image and video data. VINA uses video frames as physically grounded natural augmentations and further introduces a cross-modal supervised contrastive objective to align image and video representations under a shared real/fake decision boundary. Extensive experiments on 14 image, video, and in-the-wild benchmarks show that VINA delivers bidirectional gains, improves robustness and transferability, and achieves state-of-the-art performance across nearly all evaluated settings without complex augmentation or dataset-specific tuning.

📄 PDF Abstract BibTeX arXiv:2605.21977

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniPerson: Unified Identity-Preserving Pedestrian Generation

2025-12-02 · Changxiao Ma, Chao Yuan, Xincheng Shi, Yuzhuo Ma 외 arxiv

Person re-identification (ReID) suffers from a lack of large-scale high-quality training data due to challenges in data privacy and annotation costs. While previous approaches have explored pedestrian generation for data…

Person Re-IdentificationImage Super-ResolutionData AugmentationVideo Generation

Fine-Grained AutoAugmentation for Multi-Label Classification

2021-07-12 · Ya Wang, Hesen Chen, Fangyi Zhang, Yaohua Wang 외

Data augmentation is a commonly used approach to improving the generalization of deep learning models. Recent works show that learned data augmentation policies can achieve better generalization than hand-crafted ones. H…

ClassificationData AugmentationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+2

Self-Paced Video Data Augmentation with Dynamic Images Generated by Generative Adversarial Networks

2019-09-16 · Yumeng Zhang, Gaoguo Jia, Li Chen, Mingrui Zhang 외

There is an urgent need for an effective video classification method by means of a small number of samples. The deficiency of samples could be effectively alleviated by generating samples through Generative Adversarial N…

Data AugmentationGeneral ClassificationVideo Classification

VEnhancer: Generative Space-Time Enhancement for Video Generation

2024-07-10 · Jingwen He, Tianfan Xue, Dongyang Liu, Xinqi Lin 외

We present VEnhancer, a generative space-time enhancement framework that improves the existing text-to-video results by adding more details in spatial domain and synthetic detailed motion in temporal domain. Given a gene…

Data AugmentationSuper-ResolutionVideo GenerationVideo Super-Resolution

A Feature-space Multimodal Data Augmentation Technique for Text-video Retrieval

2022-08-03 · Alex Falcon, Giuseppe Serra, Oswald Lanz

Every hour, huge amounts of visual contents are posted on social media and user-generated content platforms. To find relevant videos by means of a natural language query, text-video retrieval methods have received increa…

Data AugmentationRetrievalVideo Retrieval