paper-with-me

Papers

Instance-aware Remote Sensing Image Captioning with Cross-hierarchy Attention

2021-05-11 · Chengze Wang, Zhiyu Jiang, Yuan Yuan

The spatial attention is a straightforward approach to enhance the performance for remote sensing image captioning. However, conventional spatial attention approaches consider only the attention distribution on one fixed coarse grid, resulting in the semantics of tiny objects can be easily ignored or disturbed during the visual feature extraction. Worse still, the fixed semantic level of conventional spatial attention limits the image understanding in different levels and perspectives, which is critical for tackling the huge diversity in remote sensing images. To address these issues, we propose a remote sensing image caption generator with instance-awareness and cross-hierarchy attention. 1) The instances awareness is achieved by introducing a multi-level feature architecture that contains the visual information of multi-level instance-possible regions and their surroundings. 2) Moreover, based on this multi-level feature extraction, a cross-hierarchy attention mechanism is proposed to prompt the decoder to dynamically focus on different semantic hierarchies and instances at each time step. The experimental results on public datasets demonstrate the superiority of proposed approach over existing methods.

📄 PDF Abstract BibTeX arXiv:2105.04996

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDiversityImage Captioning

Similar Papers 제목 키워드 기반

A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning

2025-06-11 · Swadhin Das, Divyansh Mundra, Priyanshu Dayal, Raksha Sharma

Transformer-based models have achieved strong performance in remote sensing image captioning by capturing long-range dependencies and contextual information. However, their practical deployment is hindered by high comput…

DecoderImage CaptioningKnowledge Distillation

SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

2025-07-17 · Khang Truong, Lam Pham, Hieu Tang, Jasmin Lampert 외 arxiv

Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensi…

Image Captioning

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

2026-06-15 · Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang 외 arxiv

Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared …

Representation LearningText Retrieval

Large Language Models for Captioning and Retrieving Remote Sensing Images

2024-02-09 · João Daniel Silva, João Magalhães, Devis Tuia, Bruno Martins

Image captioning and cross-modal retrieval are examples of tasks that involve the joint analysis of visual and linguistic information. In connection to remote sensing imagery, these tasks can help non-expert users in ext…

Cross-Modal RetrievalDecoderEarth ObservationImage Captioning+5

DescribeEarth: Describe Anything for Remote Sensing Images

2025-09-30 · Kaiyu Li, Zixuan Jiang, Xiangyong Cao, Jiayu Wang 외 arxiv

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing stu…

Image Captioning