paper-with-me

홈 › Papers

Direction-Oriented Visual-semantic Embedding Model for Remote Sensing Image-text Retrieval

2023-10-12 · Qing Ma, Jiancheng Pan, Cong Bai

Image-text retrieval has developed rapidly in recent years. However, it is still a challenge in remote sensing due to visual-semantic imbalance, which leads to incorrect matching of non-semantic visual and textual features. To solve this problem, we propose a novel Direction-Oriented Visual-semantic Embedding Model (DOVE) to mine the relationship between vision and language. Our highlight is to conduct visual and textual representations in latent space, directing them as close as possible to a redundancy-free regional visual representation. Concretely, a Regional-Oriented Attention Module (ROAM) adaptively adjusts the distance between the final visual and textual embeddings in the latent semantic space, oriented by regional visual features. Meanwhile, a lightweight Digging Text Genome Assistant (DTGA) is designed to expand the range of tractable textual representation and enhance global word-level semantic connections using less attention operations. Ultimately, we exploit a global visual-semantic constraint to reduce single visual dependency and serve as an external constraint for the final visual and textual representations. The effectiveness and superiority of our method are verified by extensive experiments including parameter evaluation, quantitative comparison, ablation studies and visual analysis, on two benchmark datasets, RSICD and RSITMD.

📄 PDF Abstract BibTeX arXiv:2310.08276

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalImage-text RetrievalRetrievalText Retrieval

Similar Papers 제목 키워드 기반

Remote Sensing-Oriented World Model

2025-09-22 · Yuxi Lu, Biao Wu, Zhidong Li, Kunqi Li 외 arxiv

World models have shown potential in artificial intelligence by predicting and reasoning about world states beyond direct observations. However, existing approaches are predominantly evaluated in synthetic environments o…

Spatial Reasoning

MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP

2026-01-13 · Aditya Chaudhary, Sneha Barman, Mainak Singha, Ankit Jha 외 arxiv

In this paper, we propose a novel multimodal framework, Multimodal Language-Guided Network (MMLGNet), to align heterogeneous remote sensing modalities like Hyperspectral Imaging (HSI) and LiDAR with natural language sema…

Contrastive Learning

Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation

2025-12-18 · Yunkai Yang, Yudong Zhang, Kunquan Zhang, Jinxiao Zhang 외 arxiv

With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semanti…

Semantic Segmentation

Adapting Segment Anything Model for Change Detection in HR Remote Sensing Images

2023-09-04 · Lei Ding, Kun Zhu, Daifeng Peng, Hao Tang 외

Vision Foundation Models (VFMs) such as the Segment Anything Model (SAM) allow zero-shot or interactive segmentation of visual contents, thus they are quickly applied in a variety of visual scenes. However, their direct …

Change DetectionInteractive Segmentation

Embedding Generalized Semantic Knowledge into Few-Shot Remote Sensing Segmentation

2024-05-22 · Yuyu Jia, Wei Huang, Junyu Gao, Qi Wang 외

Few-shot segmentation (FSS) for remote sensing (RS) imagery leverages supporting information from limited annotated samples to achieve query segmentation of novel classes. Previous efforts are dedicated to mining segment…

Segmentation