paper-with-me

홈 › Papers

AlignGemini: Generalizable AI-Generated Image Detection Through Task-Model Alignment

2025-12-07 · Ruoxin Chen, Jiahui Gao, Kaiqing Lin, Keyue Zhang, Yandan Zhao, Isabel Guan, Taiping Yao, Shouhong Ding arxiv

Vision Language Models (VLMs) are increasingly used for detecting AI-generated images (AIGI). However, converting VLMs into reliable detectors is resource-intensive, and the resulting models often suffer from hallucination and poor generalization. To investigate the root cause, we conduct an empirical analysis and identify two consistent behaviors. First, fine-tuning VLMs with semantic supervision improves semantic discrimination and generalizes well to unseen data. Second, fine-tuning VLMs with pixel-artifact supervision leads to weak generalization. These findings reveal a fundamental task-model misalignment. VLMs are optimized for high-level semantic reasoning and lack inductive bias toward low-level pixel artifacts. In contrast, conventional vision models effectively capture pixel-level artifacts but are less sensitive to semantic inconsistencies. This indicates that different models are naturally suited to different subtasks. Based on this insight, we formulate AIGI detection as two orthogonal subtasks: semantic consistency checking and pixel-artifact detection. Neglecting either subtask leads to systematic detection failures. We further propose the Task-Model Alignment principle and instantiate it in a two-branch detector, AlignGemini. The detector combines a VLM trained with pure semantic supervision and a vision model trained with pure pixel-artifact supervision. By enforcing clear specialization, each branch captures complementary cues. Experiments on in-the-wild benchmarks show that AlignGemini improves average accuracy by 9.5 percent using simplified training data. These results demonstrate that task-model alignment is an effective principle for generalizable AIGI detection.

📄 PDF Abstract BibTeX arXiv:2512.06746

Code (0)

등록된 구현이 없습니다.

Tasks

Artifact Detection

Similar Papers 제목 키워드 기반

Recent Advances on Generalizable Diffusion-generated Image Detection

2025-02-27 · Qijie Xu, Defang Chen, Jiawei Chen, Siwei Lyu 외

The rise of diffusion models has significantly improved the fidelity and diversity of generated images. With numerous benefits, these advancements also introduce new risks. Diffusion models can be exploited to create hig…

DiversityFace SwappingSurvey

FakeReasoning: Towards Generalizable Forgery Detection and Reasoning

2025-03-27 · Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin 외

Accurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a gener…

AttributeBinary ClassificationContrastive LearningLanguage Modeling+1

PDA: Generalizable Detection of AI-Generated Images via Post-hoc Distribution Alignment

2025-02-15 · Li Wang, Wenyu Chen, Zheng Li, Shanqing Guo

The rapid advancement of generative models has led to the proliferation of highly realistic AI-generated images, posing significant challenges for detection methods to generalize across diverse and evolving generative te…

Fake Image Detection

Generalizable AI-Generated Image Detection Based on Fractal Self-Similarity in the Spectrum

2025-03-11 · Shengpeng Xiao, Yuanfang Guo, Heqi Peng, Zeming Liu 외

The generalization performance of AI-generated image detection remains a critical challenge. Although most existing methods perform well in detecting images from generative models included in the training set, their accu…

Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection

2026-05-24 · Zijie Cao, Weijie Tu, Yao Xiao, Weijian Deng 외 arxiv

Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when gen…