paper-with-me

홈 › Papers

MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching

2025-01-20 · Yepeng Liu, Zhichao Sun, Baosheng Yu, Yitian Zhao, Bo Du, Yongchao Xu, Jun Cheng

Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The source code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2501.11299

Code (0)

등록된 구현이 없습니다.

Tasks

Keypoint DetectionZero-shot Generalization

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
BASE 설명 없음

Similar Papers 제목 키워드 기반

CAMO: Causality-Guided Adversarial Multimodal Domain Generalization for Crisis Classification

2025-12-08 · Pingchuan Ma, Chengshuai Zhao, Bohan Jiang, Saketh Vishnubhatla 외 arxiv

Crisis classification in social media aims to extract actionable disaster-related information from multimodal posts, which is a crucial task for enhancing situational awareness and facilitating timely emergency responses…

Representation LearningDomain Generalization

Adversarial Multimodal Representation Learning for Click-Through Rate Prediction

2020-03-07 · Xiang Li, Chao Wang, Jiwei Tan, Xiaoyi Zeng 외

For better user experience and business effectiveness, Click-Through Rate (CTR) prediction has been one of the most important tasks in E-commerce. Although extensive CTR prediction models have been proposed, learning goo…

Click-Through Rate PredictionPredictionRepresentation Learning

Exploiting modality-invariant feature for robust multimodal emotion recognition with missing modalities

2022-10-27 · Haolin Zuo, Rui Liu, Jinming Zhao, Guanglai Gao 외

Multimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to pre…

Emotion RecognitionMultimodal Emotion Recognition

MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis

2020-05-07 · Devamanyu Hazarika, Roger Zimmermann, Soujanya Poria

Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop soph…

Humor DetectionMultimodal Sentiment AnalysisSentiment Analysis

Self-Supervised Modality-Invariant and Modality-Specific Feature Learning for 3D Objects

2021-09-29 · Longlong Jing, Zhimin Chen, Bing Li, YingLi Tian

While most existing self-supervised 3D feature learning methods mainly focus on point cloud data, this paper explores the inherent multimodal attributes of 3D objects. We propose to jointly learn effective features from …

3D Object RecognitionCross-Modal RetrievalObject RecognitionRetrieval