paper-with-me

Papers

Large Language Models for Captioning and Retrieving Remote Sensing Images

2024-02-09 · João Daniel Silva, João Magalhães, Devis Tuia, Bruno Martins

Image captioning and cross-modal retrieval are examples of tasks that involve the joint analysis of visual and linguistic information. In connection to remote sensing imagery, these tasks can help non-expert users in extracting relevant Earth observation information for a variety of applications. Still, despite some previous efforts, the development and application of vision and language models to the remote sensing domain have been hindered by the relatively small size of the available datasets and models used in previous studies. In this work, we propose RS-CapRet, a Vision and Language method for remote sensing tasks, in particular image captioning and text-image retrieval. We specifically propose to use a highly capable large decoder language model together with image encoders adapted to remote sensing imagery through contrastive language-image pre-training. To bridge together the image encoder and language decoder, we propose training simple linear layers with examples from combining different remote sensing image captioning datasets, keeping the other parameters frozen. RS-CapRet can then generate descriptions for remote sensing images and retrieve images from textual descriptions, achieving SOTA or competitive performance with existing methods. Qualitative results illustrate that RS-CapRet can effectively leverage the pre-trained large language model to describe remote sensing images, retrieve them based on different types of queries, and also show the ability to process interleaved sequences of images and text in a dialogue manner.

📄 PDF Abstract BibTeX arXiv:2402.06475

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalDecoderEarth ObservationImage CaptioningImage RetrievalLanguage ModelingLanguage ModellingLarge Language ModelRetrieval

Similar Papers 제목 키워드 기반

Towards Automatic Satellite Images Captions Generation Using Large Language Models

2023-10-17 · Yingxu He, Qiqi Sun

Automatic image captioning is a promising technique for conveying visual information using natural language. It can benefit various tasks in satellite remote sensing, such as environmental monitoring, resource management…

Image CaptioningManagementNatural Language Understanding

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

2023-12-02 · Cong Yang, Zuchao Li, Lefei Zhang

Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this fiel…

Causal Language ModelingContrastive LearningImage CaptioningLanguage Modeling+3

Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images

2024-05-21 · Xiaofei Yu, Yitong Li, Jie Ma

Remote sensing image change captioning (RSICC) aims at generating human-like language to describe the semantic changes between bi-temporal remote sensing image pairs. It provides valuable insights into environmental dyna…

SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

2025-07-17 · Khang Truong, Lam Pham, Hieu Tang, Jasmin Lampert 외 arxiv

Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensi…

Image Captioning

Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning

2024-09-19 · Cong Yang, Zuchao Li, Hongzan Jiao, Zhi Gao 외

Recently, while significant progress has been made in remote sensing image change captioning, existing methods fail to filter out areas unrelated to actual changes, making models susceptible to irrelevant features. In th…

Change DetectionDecoderLanguage ModelingLanguage Modelling+1