paper-with-me

홈 › Papers

A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning

2024-09-27 · Swadhin Das, Raksha Sharma

Remote sensing images are highly valued for their ability to address complex real-world issues such as risk management, security, and meteorology. However, manually captioning these images is challenging and requires specialized knowledge across various domains. This letter presents an approach for automatically describing (captioning) remote sensing images. We propose a novel encoder-decoder setup that deploys a Text Graph Convolutional Network (TextGCN) and multi-layer LSTMs. The embeddings generated by TextGCN enhance the decoder's understanding by capturing the semantic relationships among words at both the sentence and corpus levels. Furthermore, we advance our approach with a comparison-based beam search method to ensure fairness in the search strategy for generating the final caption. We present an extensive evaluation of our approach against various other state-of-the-art encoder-decoder frameworks. We evaluated our method across three datasets using seven metrics: BLEU-1 to BLEU-4, METEOR, ROUGE-L, and CIDEr. The results demonstrate that our approach significantly outperforms other state-of-the-art encoder-decoder methods.

📄 PDF Abstract BibTeX arXiv:2409.18467

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderFairnessImage CaptioningSentence

Similar Papers 제목 키워드 기반

SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

2025-07-17 · Khang Truong, Lam Pham, Hieu Tang, Jasmin Lampert 외 arxiv

Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensi…

Image Captioning

Large Language Models for Captioning and Retrieving Remote Sensing Images

2024-02-09 · João Daniel Silva, João Magalhães, Devis Tuia, Bruno Martins

Image captioning and cross-modal retrieval are examples of tasks that involve the joint analysis of visual and linguistic information. In connection to remote sensing imagery, these tasks can help non-expert users in ext…

Cross-Modal RetrievalDecoderEarth ObservationImage Captioning+5

EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning

2026-08-03 · Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu 외 arxiv

Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions. Ho…

Progressive Scale-aware Network for Remote sensing Image Change Captioning

2023-03-01 · Chenyang Liu, Jiajun Yang, Zipeng Qi, Zhengxia Zou 외

Remote sensing (RS) images contain numerous objects of different scales, which poses significant challenges for the RS image change captioning (RSICC) task to identify visual changes of interest in complex scenes and des…

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

2026-04-25 · Yanpei Gong, Beichen Zhang, Hao Wang, Zhaobo Qi 외 arxiv

Remote sensing image change captioning (RSICC) aims to describe the difference between two remote sensing images. While recent methods have explored video modeling, they largely overlook the inherent ambiguities in viewp…