paper-with-me

홈 › Papers

GeoLangBind: Unifying Earth Observation with Agglomerative Vision-Language Foundation Models

2025-03-08 · Zhitong Xiong, Yi Wang, Weikang Yu, Adam J Stewart, Jie Zhao, Nils Lehmann, Thomas Dujardin, Zhenghang Yuan, Pedram Ghamisi, Xiao Xiang Zhu

Earth observation (EO) data, collected from diverse sensors with varying imaging principles, present significant challenges in creating unified analytical frameworks. We present GeoLangBind, a novel agglomerative vision--language foundation model that bridges the gap between heterogeneous EO data modalities using language as a unifying medium. Our approach aligns different EO data types into a shared language embedding space, enabling seamless integration and complementary feature learning from diverse sensor data. To achieve this, we construct a large-scale multimodal image--text dataset, GeoLangBind-2M, encompassing six data modalities. GeoLangBind leverages this dataset to develop a zero-shot foundation model capable of processing arbitrary numbers of EO data channels as input. Through our designed Modality-aware Knowledge Agglomeration (MaKA) module and progressive multimodal weight merging strategy, we create a powerful agglomerative foundation model that excels in both zero-shot vision--language comprehension and fine-grained visual understanding. Extensive evaluation across 23 datasets covering multiple tasks demonstrates GeoLangBind's superior performance and versatility in EO applications, offering a robust framework for various environmental monitoring and analysis tasks. The dataset and pretrained models will be publicly available.

📄 PDF Abstract BibTeX arXiv:2503.06312

Code (1)

xiong-zhitong/geolb-siglip 공식 구현 pytorch

Tasks

Earth Observation

Similar Papers 제목 키워드 기반

REOBench: Benchmarking Robustness of Earth Observation Foundation Models

2025-05-22 · Xiang Li, Yong Tao, Siyuan Zhang, Siwei Liu 외

Earth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce RE…

BenchmarkingContrastive LearningDisaster ResponseEarth Observation

EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answering

2023-12-19 · Junjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma 외

Earth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we dev…

ObjectObject CountingQuestion AnsweringRelational Reasoning+2

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

2026-08-13 · Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen 외 arxiv

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existin…

Spatial Reasoning

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

2024-12-19 · CVPR 2025 1 · Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz 외

Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs…

Change DetectionDisaster ResponseEarth ObservationQuestion Answering+2

OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks

2026-03-10 · Ronghao Fu, Haoran Liu, Weijie Zhang, Zhiwen Lin 외 arxiv

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth observation. However, a systematic benchm…

Contrastive LearningVisual Grounding