paper-with-me

Papers

CMIR-NET : A Deep Learning Based Model For Cross-Modal Retrieval In Remote Sensing

2019-04-09 · Ushasi Chaudhuri, Biplab Banerjee, Avik Bhattacharya, Mihai Datcu

We address the problem of cross-modal information retrieval in the domain of remote sensing. In particular, we are interested in two application scenarios: i) cross-modal retrieval between panchromatic (PAN) and multi-spectral imagery, and ii) multi-label image retrieval between very high resolution (VHR) images and speech based label annotations. Notice that these multi-modal retrieval scenarios are more challenging than the traditional uni-modal retrieval approaches given the inherent differences in distributions between the modalities. However, with the growing availability of multi-source remote sensing data and the scarcity of enough semantic annotations, the task of multi-modal retrieval has recently become extremely important. In this regard, we propose a novel deep neural network based architecture which is considered to learn a discriminative shared feature space for all the input modalities, suitable for semantically coherent information retrieval. Extensive experiments are carried out on the benchmark large-scale PAN - multi-spectral DSRSID dataset and the multi-label UC-Merced dataset. Together with the Merced dataset, we generate a corpus of speech signals corresponding to the labels. Superior performance with respect to the current state-of-the-art is observed in all the cases.

📄 PDF Abstract BibTeX arXiv:1904.04794

Code (1)

ushasi/CMIR-NET-A-deep-learning-based-model-for-cross-modal-retrieval-in-remote-sensing 공식 구현 tf

Tasks

Cross-Modal Information RetrievalCross-Modal RetrievalImage RetrievalInformation RetrievalMulti-Label Image RetrievalRetrieval

Similar Papers 제목 키워드 기반

CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval

2026-05-30 · Md Aminur Hossain, Ayush V. Patel, Nitant Dube, Biplab Banerjee arxiv

Cross-modal remote sensing image retrieval aims to retrieve semantically related scenes across heterogeneous sensing modalities. This remains challenging because paired observations may differ substantially in imaging ph…

Cross-Modal RetrievalImage Retrieval

Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization

2024-02-03 · Bo Yang, Chen Wang, Xiaoshuang Ma, Beiping Song 외

Effectively and efficiently retrieving images from remote sensing databases is a critical challenge in the realm of remote sensing big data. Utilizing hand-drawn sketches as retrieval inputs offers intuitive and user-fri…

Cross-Modal RetrievalImage RetrievalRetrievalZero-Shot Learning

AutoMIR: Effective Zero-Shot Medical Information Retrieval without Relevance Labels

2024-10-26 · Lei LI, Xiangxu Zhang, Xiao Zhou, Zheng Liu

Medical information retrieval (MIR) is essential for retrieving relevant medical knowledge from diverse sources, including electronic health records, scientific literature, and medical databases. However, achieving effec…

BenchmarkingInformation RetrievalRetrievalSelf-Learning

Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval

2024-11-22 · Zengbao Sun, Ming Zhao, Gaorui Liu, André Kaup

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in…

Image RetrievalRerankingRetrievalText Retrieval+1

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

2025-12-12 · Emanuel Sánchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu 외 arxiv

Satellite imagery differs fundamentally from natural images: its aerial viewpoint, very high resolution, diverse scale variations, and abundance of small objects demand both region-level spatial reasoning and holistic sc…

Cross-Modal RetrievalScene ClassificationScene UnderstandingQuestion Answering