paper-with-me

홈 › Papers

FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images

2025-07-01 · Guang Yang arxiv

The rapid advancement of diffusion models, particularly Stable Diffusion 3.5, has enabled the generation of highly photorealistic synthetic images that pose significant challenges to existing detection methods. This paper presents FreqCross, a novel multi-modal fusion network that combines spatial RGB features, frequency domain artifacts, and radial energy distribution patterns to achieve robust detection of AI-generated images. Our approach leverages a three-branch architecture: (1) a ResNet-18 backbone for spatial feature extraction, (2) a lightweight CNN for processing 2D FFT magnitude spectra, and (3) a multi-layer perceptron for analyzing radial energy profiles. We introduce a novel radial energy distribution analysis that captures characteristic frequency artifacts inherent in diffusion-generated images, and fuse it with spatial and spectral cues via simple feature concatenation followed by a compact classification head. Extensive experiments on a dataset of 10,000 paired real (MS-COCO) and synthetic (Stable Diffusion 3.5) images demonstrate that FreqCross achieves 97.8\% accuracy, outperforming state-of-the-art baselines by 5.2\%. The frequency analysis further reveals that synthetic images exhibit distinct spectral signatures in the 0.1--0.4 normalised frequency range, providing theoretical foundation for our approach. Code and pre-trained models are publicly available to facilitate reproducible research.

📄 PDF Abstract BibTeX arXiv:2507.02995

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion

2026-02-04 · Yixin Zhu, Long Lv, Pingping Zhang, Xuehu Liu 외 arxiv

Multi-Modal Image Fusion (MMIF) aims to combine images from different modalities to produce fused images, retaining texture details and preserving significant information. Recently, some MMIF methods incorporate frequenc…

Frequency Spectrum is More Effective for Multimodal Representation and Fusion: A Multimodal Spectrum Rumor Detector

2023-12-18 · An Lao, Qi Zhang, Chongyang Shi, Longbing Cao 외

Multimodal content, such as mixing text with images, presents significant challenges to rumor detection in social media. Existing multimodal rumor detection has focused on mixing tokens among spatial and sequential locat…

Contrastive Learning

A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification

2025-10-06 · Hao Liu, Yunhao Gao, Wei Li, Mingyang Zhang 외 arxiv

Deep learning-based methods have achieved significant success in remote sensing Earth observation data analysis. Numerous feature fusion techniques address multimodal remote sensing image classification by integrating gl…

Remote Sensing Image Classification

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

2025-08-21 · Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang 외 arxiv

Multimodal Image Fusion (MMIF) aims to integrate complementary information from different imaging modalities to overcome the limitations of individual sensors. It enhances image quality and facilitates downstream applica…

AdaFuse: Adaptive Medical Image Fusion Based on Spatial-Frequential Cross Attention

2023-10-09 · Xianming Gu, Lihui Wang, Zeyu Deng, Ying Cao 외

Multi-modal medical image fusion is essential for the precise clinical diagnosis and surgical navigation since it can merge the complementary information in multi-modalities into a single image. The quality of the fused …