paper-with-me

Papers

LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival

2024-03-16 · Yuanxin Zhao, Mi Zhang, Bingnan Yang, Zhan Zhang, Jiaju Kang, Jianya Gong

Image-text retrieval (ITR) plays a significant role in making informed decisions for various remote sensing (RS) applications. Nonetheless, creating ITR datasets containing vision and language modalities not only requires significant geo-spatial sampling area but also varing categories and detailed descriptions. To this end, we introduce an image caption dataset LuojiaHOG, which is geospatial-aware, label-extension-friendly and comprehensive-captioned. LuojiaHOG involves the hierarchical spatial sampling, extensible classification system to Open Geospatial Consortium (OGC) standards, and detailed caption generation. In addition, we propose a CLIP-based Image Semantic Enhancement Network (CISEN) to promote sophisticated ITR. CISEN consists of two components, namely dual-path knowledge transfer and progressive cross-modal feature fusion. Comprehensive statistics on LuojiaHOG reveal the richness in sampling diversity, labels quantity and descriptions granularity. The evaluation on LuojiaHOG is conducted across various state-of-the-art ITR models, including ALBEF, ALIGN, CLIP, FILIP, Wukong, GeoRSCLIP and CISEN. We use second- and third-level labels to evaluate these vision-language models through adapter-tuning and CISEN demonstrates superior performance. For instance, it achieves the highest scores with WMAP@5 of 88.47\% and 87.28\% on third-level ITR tasks, respectively. In particular, CISEN exhibits an improvement of approximately 1.3\% and 0.9\% in terms of WMAP@5 compared to its baseline. These findings highlight CISEN advancements accurately retrieving pertinent information across image and text. LuojiaHOG and CISEN can serve as a foundational resource for future RS image-text alignment research, facilitating a wide range of vision-language applications.

📄 PDF Abstract BibTeX arXiv:2403.10887

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationImage-text RetrievalText RetrievalTransfer Learning

Methods 이 논문이 사용한 방법론

ALBEF ALBEF introduces a contrastive loss to align the image and text representations before fusing them through cross-modal attention. This enables more grounded vision and language…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Instance-aware Remote Sensing Image Captioning with Cross-hierarchy Attention

2021-05-11 · Chengze Wang, Zhiyu Jiang, Yuan Yuan

The spatial attention is a straightforward approach to enhance the performance for remote sensing image captioning. However, conventional spatial attention approaches consider only the attention distribution on one fixed…

DecoderDiversityImage Captioning

LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences

2025-07-25 · Yusuke Hirota, Boyi Li, Ryo Hachiuma, Yueh-Hua Wu 외 arxiv

Large Vision-Language Models (LVLMs) have transformed image captioning, shifting from concise captions to detailed descriptions. We introduce LOTUS, a leaderboard for evaluating detailed captions, addressing three main g…

Image Captioning

Uncertainty-Aware Image Captioning

2022-11-30 · Zhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 외

It is well believed that the higher uncertainty in a word of the caption, the more inter-correlated context information is required to determine it. However, current image captioning methods usually consider the generati…

Caption GenerationImage CaptioningSentence

Phrase-based Image Captioning with Hierarchical LSTM Model

2017-11-11 · Ying Hua Tan, Chee Seng Chan

Automatic generation of caption to describe the content of an image has been gaining a lot of research interests recently, where most of the existing works treat the image caption as pure sequential data. Natural languag…

DecoderImage CaptioningImage Descriptionmodel+1

Object-oriented backdoor attack against image captioning

2024-01-05 · Meiling Li, Nan Zhong, Xinpeng Zhang, Zhenxing Qian 외

Backdoor attack against image classification task has been widely studied and proven to be successful, while there exist little research on the backdoor attack against vision-language models. In this paper, we explore ba…

Backdoor AttackImage Captioningimage-classificationImage Classification+1