paper-with-me

홈 › Papers

OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data

2025-05-29 · Fengxiang Wang, Mingshuo Chen, Xuming He, Yifan Zhang, Feng Liu, Zijie Guo, Zhenghao Hu, Jiong Wang, Jingyi Xu, Zhangrui Li, Fenghua Ling, Ben Fei, Weijia Li, Long Lan, Wenjing Yang, Wenlong Zhang, Lei Bai

Existing benchmarks for Earth science multimodal learning exhibit critical limitations in systematic coverage of geosystem components and cross-sphere interactions, often constrained to isolated subsystems (only in Human-activities sphere or atmosphere) with limited evaluation dimensions (less than 16 tasks). To address these gaps, we introduce OmniEarth-Bench, the first comprehensive multimodal benchmark spanning all six Earth science spheres (atmosphere, lithosphere, Oceansphere, cryosphere, biosphere and Human-activities sphere) and cross-spheres with one hundred expert-curated evaluation dimensions. Leveraging observational data from satellite sensors and in-situ measurements, OmniEarth-Bench integrates 29,779 annotations across four tiers: perception, general reasoning, scientific knowledge reasoning and chain-of-thought (CoT) reasoning. This involves the efforts of 2-5 experts per sphere to establish authoritative evaluation dimensions and curate relevant observational datasets, 40 crowd-sourcing annotators to assist experts for annotations, and finally, OmniEarth-Bench is validated via hybrid expert-crowd workflows to reduce label ambiguity. Experiments on 9 state-of-the-art MLLMs reveal that even the most advanced models struggle with our benchmarks, where none of them reach 35\% accuracy. Especially, in some cross-spheres tasks, the performance of leading models like GPT-4o drops to 0.0\%. OmniEarth-Bench sets a new standard for geosystem-aware AI, advancing both scientific discovery and practical applications in environmental monitoring and disaster prediction. The dataset, source code, and trained models were released.

📄 PDF Abstract BibTeX arXiv:2505.23522

Code (0)

등록된 구현이 없습니다.

Tasks

scientific discovery

Similar Papers 제목 키워드 기반

OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks

2026-03-10 · Ronghao Fu, Haoran Liu, Weijie Zhang, Zhiwen Lin 외 arxiv

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth observation. However, a systematic benchm…

Contrastive LearningVisual Grounding

EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models

2025-05-22 · Wanghan Xu, Xiangyu Zhao, Yuhao Zhou, Xiaoyu Yue 외

Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Ear…

Question AnsweringSpecificity

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs

2026-05-08 · Song Zhang, Yanlong Chen, Yilin Li, Yining Chen 외 arxiv

Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radically different visual evidence across ground sampling distances (GSDs)…

parameter-efficient fine-tuningAnswer Generation

Toward Artificial Intelligence Enabled Earth System Coupling

2026-03-26 · Maria Kaselimi, Anna Belehaki arxiv

Coupling constitutes a foundational mechanism in the Earth system, regulating the interconnected physical, chemical, and biological processes that link its spheres. This review examines how emerging artificial intelligen…

A Self-Evolving AI Agent System for Climate Science

2025-07-23 · Zijie Guo, Jiong Wang, Fenghua Ling, Wangxu Wei 외 arxiv

Scientific progress in Earth science depends on integrating data across the planet's interconnected spheres. However, the accelerating volume and fragmentation of multi-sphere knowledge and data have surpassed human anal…