paper-with-me

Papers

Exploring Models and Data for Remote Sensing Image Caption Generation

2017-12-21 · Xiaoqiang Lu, Binqiang Wang, Xiangtao Zheng, Xuelong. Li

Inspired by recent development of artificial satellite, remote sensing images have attracted extensive attention. Recently, noticeable progress has been made in scene classification and target detection.However, it is still not clear how to describe the remote sensing image content with accurate and concise sentences. In this paper, we investigate to describe the remote sensing images with accurate and flexible sentences. First, some annotated instructions are presented to better describe the remote sensing images considering the special characteristics of remote sensing images. Second, in order to exhaustively exploit the contents of remote sensing images, a large-scale aerial image data set is constructed for remote sensing image caption. Finally, a comprehensive review is presented on the proposed data set to fully advance the task of remote sensing caption. Extensive experiments on the proposed data set demonstrate that the content of the remote sensing image can be completely described by generating language descriptions. The data set is available at https://github.com/201528014227051/RSICD_optimal

📄 PDF Abstract BibTeX arXiv:1712.07835

Code (2)

201528014227051/RSICD_optimal 공식 구현
arampacha/clip-rsicd jax

Tasks

Caption GenerationImage-to-Text RetrievalScene Classification

Similar Papers 제목 키워드 기반

VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

2024-06-18 · Xiang Li, Jian Ding, Mohamed Elhoseiny

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been pro…

Image CaptioningQuestion AnsweringVisual GroundingVisual Question Answering

Multilingual Vision-Language Pre-training for the Remote Sensing Domain

2024-10-30 · João Daniel Silva, Joao Magalhaes, Devis Tuia, Bruno Martins

Methods based on Contrastive Language-Image Pre-training (CLIP) are nowadays extensively used in support of vision-and-language tasks involving remote sensing data, such as cross-modal retrieval. The adaptation of CLIP t…

Cross-Modal Retrievalimage-classificationImage ClassificationImage-text Retrieval+4

Changes to Captions: An Attentive Network for Remote Sensing Change Captioning

2023-04-03 · Shizhen Chang, Pedram Ghamisi

In recent years, advanced research has focused on the direct learning and analysis of remote sensing images using natural language processing (NLP) techniques. The ability to accurately describe changes occurring in mult…

SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

2025-07-17 · Khang Truong, Lam Pham, Hieu Tang, Jasmin Lampert 외 arxiv

Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensi…

Image Captioning

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

2023-12-02 · Cong Yang, Zuchao Li, Lefei Zhang

Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this fiel…

Causal Language ModelingContrastive LearningImage CaptioningLanguage Modeling+3