paper-with-me

Papers

Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonization

2025-01-01 · CVPR 2025 1 · Kai Mao, Ping Wei, Yiyang Lian, Yangyang Wang, Nanning Zheng

Anomaly detection is a significant task for its application and research value. While existing methods have made impressive progress within the same modality, cross-modal anomaly detection remains an open and challenging problem. In this paper, we propose a cross-modal anomaly detection model that is trained using data from a variety of existing modalities and can be generalized well to unseen modalities. The model consists of three major components: 1) the Transferable Visual Prototype directly learns normal/abnormal semantics in visual space; 2) the Prototype Harmonization strategy adaptively utilizes the Transferable Visual Prototypes from various modalities for inference on the unknown modality; 3) the Visual Discrepancy Inference under the few-shot setting enhances performance. In the zero-shot setting, the proposed method achieves AUROC improvements of 4.1%, 6.1%, 7.6%, and 6.8% over the best competing methods in the RGB, 3D, MRI/CT, and Thermal modalities, respectively. In the few-shot setting, our model also achieves the highest AUROC/AP on ten datasets in four modalities, substantially outperforming existing methods. Codes are available at https://github.com/Kerio99/CMAD.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Anomaly Detection

Similar Papers 제목 키워드 기반

EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration

2026-07-05 · Shen Shen, Quan Zhang, Dan Jiang, Ke Zhang arxiv

The rapid proliferation of artificial intelligence-generated content necessitates reliable multimodal forensics. Beyond video-level binary classification, precisely localizing sparsely distributed forged segments in long…

Representation LearningBinary Classification

Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics

2025-12-03 · Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng, Jie Wang 외 arxiv

Video frame interpolation has long been challenged by limited controllability and interactivity, especially in scenarios involving fast, highly non-linear, and fine-grained motion. Although recent interactive interpolati…

Video Frame Interpolation

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

2026-05-22 · Jiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang 외 arxiv

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degr…

Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

2026-07-16 · Xiao Wang, Bing Wang, Bin Yang, Cuiqun Chen 외 arxiv

Person re-identification (ReID) serves as a critical component in intelligent surveillance systems, aiming to match identities across disjoint camera networks. While traditional methods primarily rely on single-modal RGB…

Person Re-Identification

Defending Multimodal Fusion Models against Single-Source Adversaries

2022-06-25 · CVPR 2021 1 · Karren Yang, Wan-Yi Lin, Manash Barman, Filipe Condessa 외

Beyond achieving high performance across many vision tasks, multimodal models are expected to be robust to single-source faults due to the availability of redundant information between modalities. In this paper, we inves…

Action Recognitionobject-detectionObject DetectionSentiment Analysis