paper-with-me

Papers

A Prior Instruction Representation Framework for Remote Sensing Image-text Retrieval

2023-10-27 · ACMMM 2023 10 · Jiancheng Pan, Qing Ma, Cong Bai

This paper presents a prior instruction representation framework (PIR) for remote sensing image-text retrieval, aimed at remote sensing vision-language understanding tasks to solve the semantic noise problem. Our highlight is the proposal of a paradigm that draws on prior knowledge to instruct adaptive learning of vision and text representations. Concretely, two progressive attention encoder (PAE) structures, Spatial-PAE and Temporal-PAE, are proposed to perform long-range dependency modeling to enhance key feature representation. In vision representation, Vision Instruction Representation (VIR) based on Spatial-PAE exploits the prior-guided knowledge of the remote sensing scene recognition by building a belief matrix to select key features for reducing the impact of semantic noise. In text representation, Language Cycle Attention (LCA) based on Temporal-PAE uses the previous time step to cyclically activate the current time step to enhance text representation capability. A cluster-wise affiliation loss is proposed to constrain the inter-classes and to reduce the semantic confusion zones in the common subspace. Comprehensive experiments demonstrate that using prior knowledge instruction could enhance vision and text representations and could outperform the state-of-the-art methods on two benchmark datasets, RSICD and RSITMD.

📄 PDF Abstract BibTeX

Code (1)

jaychempan/PIR 공식 구현 pytorch

Tasks

Cross-Modal RetrievalImage-text RetrievalRetrievalScene RecognitionText Retrieval

Similar Papers 제목 키워드 기반

PIR: Remote Sensing Image-Text Retrieval with Prior Instruction Representation Learning

2024-05-16 · Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai 외

Remote sensing image-text retrieval constitutes a foundational aspect of remote sensing interpretation tasks, facilitating the alignment of vision and language representations. This paper introduces a prior instruction r…

Image-text RetrievalRepresentation LearningRetrievalScene Recognition+1

Falcon: A Remote Sensing Vision-Language Foundation Model

2025-03-14 · Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu 외

This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensi…

Image Captioningimage-classificationImage Classificationobject-detection+1

Remote Sensing-Oriented World Model

2025-09-22 · Yuxi Lu, Biao Wu, Zhidong Li, Kunqi Li 외 arxiv

World models have shown potential in artificial intelligence by predicting and reasoning about world states beyond direct observations. However, existing approaches are predominantly evaluated in synthetic environments o…

Spatial Reasoning

DDFAV: Remote Sensing Large Vision Language Models Dataset and Evaluation Benchmark

2024-11-05 · Haodong Li, Haicheng Qu, Xiaofeng Zhang

With the rapid development of large vision language models (LVLMs), these models have shown excellent results in various multimodal tasks. Since LVLMs are prone to hallucinations and there are currently few datasets and …

Data AugmentationHallucinationHallucination Evaluation

Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning

2024-09-19 · Cong Yang, Zuchao Li, Hongzan Jiao, Zhi Gao 외

Recently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models susceptible to irrelevant features. In th…

Change DetectionDecoderLanguage ModelingLanguage Modelling+1