paper-with-me

홈 › Papers

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

2026-06-09 · Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li arxiv

RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery. However, existing models support only a narrow range of sensor types and tasks, yielding a fragmented view of the earth and leaving cross-modal geoscientific knowledge largely unexploited. This work presents Earth-OneVision, a 2B RS-MLLM that unifies six sensor modalities (i.e., optical, SAR, infrared, multispectral, temporal, and video) and cross-sensor fusion across 9 task categories within a single autoregressive framework. Three dedicated mechanisms address three bottlenecks. Full-Granularity Vision-Language Alignment (FGVLA) aligns multi-level visual features with the multi-dimensional language space. Spatial-Linguistic Isomorphic Serialization (SLIS) unifies heterogeneous spatial outputs as autoregressive tokens. Progressive Cross-Modality Adaptation (PCMA) decomposes the compound domain gap into sequential stages, tackling the viewpoint and imaging physics gaps in turn. To support joint training, MMRS-OneVision is constructed with ~34M QA pairs spanning all six sensor modalities and cross-sensor fusion across 9 task categories, substantially exceeding existing RS multimodal instruction datasets. With only 2B parameters, Earth-OneVision achieves competitive or state-of-the-art results across extensive benchmarks, consistently matching or outperforming 4B-72B RS-MLLMs. It achieves 87.52% P@0.5 on the OPT-RSVG testset for optical visual grounding and 80.68% on the SAR VQA benchmark SARLANG-Bench, exceeding 7B models by over 7%. It further achieves 75.74% recall on the BigEarthNet-MS testset for multispectral classification, and 81.94% MCQ accuracy on EarthMind-Bench for cross-modality reasoning.

📄 PDF Abstract BibTeX arXiv:2606.10819

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

MultiEarth 2023 -- Multimodal Learning for Earth and Environment Workshop and Challenge

2023-06-07 · Miriam Cha, Gregory Angelides, Mark Hamilton, Andy Soszynski 외

The Multimodal Learning for Earth and Environment Workshop (MultiEarth 2023) is the second annual CVPR workshop aimed at the monitoring and analysis of the health of Earth ecosystems by leveraging the vast amount of remo…

Representation Learning

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

2026-04-19 · Rui Min, Liang Yao, Shiyu Miao, Shengxiang Xu 외 arxiv

A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Remote Sensing MLLMs fail to meet this requ…

Advances and Challenges in Multimodal Remote Sensing Image Registration

2023-02-02 · Bai Zhu, Liang Zhou, Simiao Pu, Jianwei Fan 외

Over the past few decades, with the rapid development of global aerospace and aerial remote sensing technology, the types of sensors have evolved from the traditional monomodal sensors (e.g., optical sensors) to the new …

Image Registration

Galileo: Learning Global and Local Features in Pretrained Remote Sensing Models

2025-02-13 · Gabriel Tseng, Anthony Fuller, Marlena Reil, Henry Herzog 외

From crop mapping to flood detection, machine learning in remote sensing has a wide range of societally beneficial applications. The commonalities between remote sensing data in these applications present an opportunity …

Self-Supervised Learning

MANet: Fine-Tuning Segment Anything Model for Multimodal Remote Sensing Semantic Segmentation

2024-10-15 · Xianping Ma, Xiaokang Zhang, Man-on Pun, Bo Huang

Multimodal remote sensing data, collected from a variety of sensors, provide a comprehensive and integrated perspective of the Earth's surface. By employing multimodal fusion techniques, semantic segmentation offers more…

General KnowledgeSegmentationSemantic Segmentation