paper-with-me

홈 › Papers

UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

2025-09-30 · Yuan Zhao, Youwei Pang, Lihe Zhang, Hanqi Liu, Jiaming Zuo, Huchuan Lu, Xiaoqi Zhao arxiv

Existing anomaly detection (AD) methods often treat the modality and class as independent factors. Although this paradigm has enriched the development of AD research branches and produced many specialized models, it has also led to fragmented solutions and excessive memory overhead. Moreover, reconstruction-based multi-class approaches typically rely on shared decoding paths, which struggle to handle large variations across domains, resulting in distorted normality boundaries, domain interference, and high false alarm rates. To address these limitations, we propose UniMMAD, a unified framework for multi-modal and multi-class anomaly detection. At the core of UniMMAD is a Mixture-of-Experts (MoE)-driven feature decompression mechanism, which enables adaptive and disentangled reconstruction tailored to specific domains. This process is guided by a ``general to specific'' paradigm. In the encoding stage, multi-modal inputs of varying combinations are compressed into compact, general-purpose features. The encoder incorporates a feature compression module to suppress latent anomalies, encourage cross-modal interaction, and avoid shortcut learning. In the decoding stage, the general features are decompressed into modality-specific and class-specific forms via a sparsely-gated cross MoE, which dynamically selects expert pathways based on input modality and class. To further improve efficiency, we design a grouped dynamic filtering mechanism and a MoE-in-MoE structure, reducing parameter usage by 75\% while maintaining sparse activation and fast inference. UniMMAD achieves state-of-the-art performance on 9 anomaly detection datasets, spanning 3 fields, 12 modalities, and 66 classes. The source code will be available at https://github.com/yuanzhao-CVLAB/UniMMAD.

📄 PDF Abstract BibTeX arXiv:2509.25934

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-class Anomaly Detection

Similar Papers 제목 키워드 기반

Text-Guided Multimodal Unified Industrial Anomaly Detection

2026-04-24 · Zewen Li, Shuo Ye, Zitong Yu, Weicheng Xie 외 arxiv

Industrial anomaly detection based on RGB-3D multimodal data has emerged as a mainstream paradigm for intelligent quality inspection. However, existing unsupervised methods suffer from two critical limitations: ambiguous…

Anomaly Detection

Uni-RCM: Unified Reference-guided Cross-modal Mapping for Multi-Class Anomaly Detection

2026-05-28 · Yangchen Wu, Huiqiang Xie arxiv

Multi-modal industrial anomaly detection typically relies on separate models for each product category, fundamentally limiting practical scalability. When shifting to a unified paradigm that handles diverse classes simul…

Multi-class Anomaly Detection

Open-set Cross Modal Generalization via Multimodal Unified Representation

2025-07-20 · Hai Huang, Yan Xia, Shulei Wang, Hanting Wang 외 arxiv

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in o…

Self-Supervised LearningContrastive Learning

One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection

2025-10-20 · Jia Guo, Shuai Lu, Lei Fan, Zelin Li 외 arxiv

Unsupervised anomaly detection (UAD) has evolved from building specialized single-class models to unified multi-class models, yet existing multi-class models significantly underperform the most advanced one-for-one count…

Unsupervised Anomaly Detection

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

2025-05-08 · Chao Liao, Liyang Liu, Xun Wang, Zhengxiong Luo 외

Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple modalities. In this paper, we present Mo…

Image GenerationText GenerationText to Image GenerationText-to-Image Generation