paper-with-me

Papers

ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models

2024-02-17 · Zhenghang Yuan, Zhitong Xiong, Lichao Mou, Xiao Xiang Zhu

An in-depth comprehension of global land cover is essential in Earth observation, forming the foundation for a multitude of applications. Although remote sensing technology has advanced rapidly, leading to a proliferation of satellite imagery, the inherent complexity of these images often makes them difficult for non-expert users to understand. Natural language, as a carrier of human knowledge, can be a bridge between common users and complicated satellite imagery. In this context, we introduce a global-scale, high-quality image-text dataset for remote sensing, providing natural language descriptions for Sentinel-2 data to facilitate the understanding of satellite imagery for common users. Specifically, we utilize Sentinel-2 data for its global coverage as the foundational image source, employing semantic segmentation labels from the European Space Agency's (ESA) WorldCover project to enrich the descriptions of land covers. By conducting in-depth semantic analysis, we formulate detailed prompts to elicit rich descriptions from ChatGPT. To enhance the dataset's quality, we introduce the manual verification process. This step involves manual inspection and correction to refine the dataset, thus significantly improving its accuracy and quality. Finally, we offer the community ChatEarthNet, a large-scale image-text dataset characterized by global coverage, high quality, wide-ranging diversity, and detailed descriptions. ChatEarthNet consists of 163,488 image-text pairs with captions generated by ChatGPT-3.5 and an additional 10,000 image-text pairs with captions generated by ChatGPT-4V(ision). This dataset has significant potential for training vision-language geo-foundation models and evaluating large vision-language models for remote sensing. The dataset will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2402.11325

Code (1)

zhu-xlab/ChatEarthNet 공식 구현

Tasks

Earth ObservationImage CaptioningSemantic Segmentation

Similar Papers 제목 키워드 기반

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model

2025-01-01 · Chenyang Liu, Keyan Chen, Rui Zhao, Zhengxia Zou 외

Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lac…

Image Generation

LM-Net: A Light-weight and Multi-scale Network for Medical Image Segmentation

2025-01-07 · Zhenkun Lu, Chaoyin She, Wei Wang, Qinghua Huang

Current medical image segmentation approaches have limitations in deeply exploring multi-scale information and effectively combining local detail textures with global contextual semantic information. This results in over…

Image SegmentationMedical Image SegmentationSegmentationSemantic Segmentation

Multi-scale gridded Gabor attention for cirrus segmentation

2024-07-11 · Felix Richards, Adeline Paiement, Xianghua Xie, Elisabeth Sola 외

In this paper, we address the challenge of segmenting global contaminants in large images. The precise delineation of such structures requires ample global context alongside understanding of textural patterns. CNNs speci…

Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation

2026-02-25 · Junxin Lu, Tengfei Song, Zhanglin Wu, Pengfei Li 외 arxiv

Text Image Machine Translation (TIMT) aims to translate text embedded in images in the source-language into target-language, requiring synergistic integration of visual perception and linguistic understanding. Existing T…

Machine Translation

Asymmetric Cross-Scale Alignment for Text-Based Person Search

2022-11-26 · Zhong Ji, Junhua Hu, Deyin Liu, Lin Yuanbo Wu 외

Text-based person search (TBPS) is of significant importance in intelligent surveillance, which aims to retrieve pedestrian images with high semantic relevance to a given text description. This retrieval task is characte…

cross-modal alignmentPerson SearchRetrievalSentence+1