paper-with-me

홈 › Papers

EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models

2025-05-22 · Wanghan Xu, Xiangyu Zhao, Yuhao Zhou, Xiaoyu Yue, Ben Fei, Fenghua Ling, Wenlong Zhang, Lei Bai

Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated subdomains, lacking holistic evaluation. Furthermore, current benchmarks typically neglect the assessment of LLMs' capabilities in open-ended scientific exploration. In this paper, we present a comprehensive and professional benchmark for the Earth sciences, designed to evaluate the capabilities of LLMs in scientific exploration within this domain, spanning from fundamental to advanced levels. Leveraging a corpus of 100,000 research papers, we first construct two Question Answering (QA) datasets: Earth-Iron, which offers extensive question coverage for broad assessment, and Earth-Silver, which features a higher level of difficulty to evaluate professional depth. These datasets encompass five Earth spheres, 114 disciplines, and 11 task categories, assessing foundational knowledge crucial for scientific exploration. Most notably, we introduce Earth-Gold with new metrics, a dataset comprising open-ended multi-turn dialogues specifically designed to evaluate the advanced capabilities of LLMs in scientific exploration, including methodology induction, limitation analysis, and concept proposal. Extensive experiments reveal limitations in 11 leading LLMs across different domains and tasks, highlighting considerable room for improvement in their scientific exploration capabilities. The benchmark is available on https://huggingface.co/ai-earth .

📄 PDF Abstract BibTeX arXiv:2505.17139

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSpecificity

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

OpenEarthSensing: Large-Scale Fine-Grained Benchmark for Open-World Remote Sensing

2025-02-28 · Xiang Xiang, Zhuo Xu, Yao Deng, Qinhao Zhou 외

In open-world remote sensing, deployed models must continuously adapt to a steady influx of new data, which often exhibits various shifts compared to what the model encountered during the training phase. To effectively h…

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning

2026-05-11 · Wanghan Xu, Yuhao Zhou, Hengyuan Zhao, Shuo Li 외 arxiv

Large language models can fail in critic interaction not only by answering incorrectly, but also by abandoning an initially correct scientific solution after user criticism. This is especially risky in scientific reasoni…

Reinforcement Learning

GeoR-Bench: Evaluating Geoscience Visual Reasoning

2026-05-12 · Yushuo Zheng, Zicheng Zhang, Huiyu Duan, Chunyi Li 외 arxiv

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental prote…

Visual Reasoning

PLUME: Procedural Layer Underground Modeling Engine

2025-08-28 · Gabriel Manuel Garcia, Antoine Richard, Miguel Olivares-Mendez arxiv

As space exploration advances, underground environments are becoming increasingly attractive due to their potential to provide shelter, easier access to resources, and enhanced scientific opportunities. Although such env…

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

2026-09-24 · Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen 외 hf

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: …