paper-with-me

홈 › Papers

Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection

2026-04-14 · Haifeng Zhang, Qinghui He, Xiuli Bi, Bo Liu, Chi-Man Pun, Bin Xiao arxiv

In recent years, the rapid development of generative artificial intelligence technology has significantly lowered the barrier to creating high-quality fake images, posing a serious challenge to information authenticity and credibility. Existing generated image detection methods typically enhance generalization through model architecture or network design. However, their generalization performance remains susceptible to data bias, as the training data may drive models to fit specific generative patterns and content rather than the common features shared by images from different generative models (asymmetric bias learning). To address this issue, we propose a Multi-dimensional Adversarial Feature Learning (MAFL) framework. The framework adopts a pretrained multimodal image encoder as the feature extraction backbone, constructs a real-fake feature learning network, and designs an adversarial bias-learning branch equipped with a multi-dimensional adversarial loss, forming an adversarial training mechanism between authenticity-discriminative feature learning and bias feature learning. By suppressing generation-pattern and content biases, MAFL guides the model to focus on the generative features shared across different generative models, thereby effectively capturing the fundamental differences between real and generated images, enhancing cross-model generalization, and substantially reducing the reliance on large-scale training data. Through extensive experimental validation, our method outperforms existing state-of-the-art approaches by 10.89% in accuracy and 8.57% in Average Precision (AP). Notably, even when trained with only 320 images, it can still achieve over 80% detection accuracy on public datasets.

📄 PDF Abstract BibTeX arXiv:2604.12353

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language Inference

2020-10-08 · CONLL 2020 · Tianyu Liu, Xin Zheng, Xiaoan Ding, Baobao Chang 외

The prior work on natural language inference (NLI) debiasing mainly targets at one or few known biases while not necessarily making the models more robust. In this paper, we focus on the model-agnostic debiasing strategi…

Data AugmentationMixture-of-ExpertsNatural Language Inference

It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations

2020-05-09 · ACL 2020 6 · Samson Tan, Shafiq Joty, Min-Yen Kan, Richard Socher

Training on only perfect Standard English corpora predisposes pre-trained neural networks to discriminate against minorities from non-standard linguistic backgrounds (e.g., African American Vernacular English, Colloquial…

A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa

2025-11-21 · Tim Schlippe, Matthias Wölfel, Koena Ronny Mabokela arxiv

This study investigates how cultural proximity affects the ability to detect AI-generated fake news by comparing South African participants with those from other nationalities. As large language models increasingly enabl…

Combating Adversarial Attacks with Multi-Agent Debate

2024-01-11 · Steffi Chern, Zhen Fan, Andy Liu

While state-of-the-art language models have achieved impressive results, they remain susceptible to inference-time adversarial attacks, such as adversarial prompts generated by red teams arXiv:2209.07858. One approach pr…

Language ModelingLanguage Modelling

Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization

2024-12-15 · Portia Cooper, Harshita Narnoli, Mihai Surdeanu

Text-to-image models are vulnerable to the stepwise "Divide-and-Conquer Attack" (DACA) that utilize a large language model to obfuscate inappropriate content in prompts by wrapping sensitive text in a benign narrative. T…

Adversarial TextBinary ClassificationLanguage ModelingLanguage Modelling+2