paper-with-me

홈 › Papers

Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image Forensics

2026-05-13 · Chiara Musso, Joy Battocchio, Andrea Montibeller, Giulia Boato arxiv

As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. However, the prevailing reliance on frozen, pre-trained backbones introduces a subtle yet critical vulnerability. In this work, we present the Surrogate Iterative Adversarial Attack (SIAA), a gray-box attack that exploits knowledge of the detector's ViT backbone alone and operates entirely within the target detector's feature space to craft highly effective adversarial examples. Through our experiments, involving multiple ViT-based detectors and diverse gray-box scenarios, including few-shot learning, complete training misalignment and attack transferability tests, we demonstrate that this vulnerability consistently yields high attack success rates, often approaching white-box performance. By doing so, we reveal that backbone knowledge alone is sufficient to undermine detector reliability, highlighting the urgent need for more resilient defenses in adversarial multimedia forensics.

📄 PDF Abstract BibTeX arXiv:2605.13381

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial AttackDeepFake DetectionFew-Shot Learning

Similar Papers 제목 키워드 기반

Align-RAG: Alignment Is All You Need for TSFM In-Context Learning

2026-08-06 · Mohammad Asadi, Soheil Hor, Bardiya Akhbari, Jack W. O'Sullivan 외 arxiv

Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that …

Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models

2024-10-25 · Shenghao Fu, Junkai Yan, Qize Yang, Xihan Wei 외

Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning t…

object-detectionObject Detection

SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation

2026-01-25 · Taewan Cho, Taeryang Kim, Andrew Jaeyong Choi arxiv

Robotic and autonomous systems need dense spatial cues, but many monocular depth models are heavy, task-specific, or hard to attach to an existing multimodal stack. CLIP offers strong semantic representations, yet most C…

Monocular Depth Estimation

Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

2026-07-11 · Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia 외 arxiv

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear. We benchmark 15 foundation-model backbones across breast density, BI-…

Few-Shot Continual Learning for 3D Brain MRI with Frozen Foundation Models

2026-02-26 · Chi-Sheng Chen, Xinyu Zhang, Guan-Ying Chen, Qiuzhe Xie 외 arxiv

Foundation models pretrained on large-scale 3D medical imaging data face challenges when adapted to multiple downstream tasks under continual learning with limited labeled data. We address few-shot continual learning for…

Continual LearningTumor SegmentationAge Estimation