paper-with-me

Papers

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?

2026-06-11 · Dat Tien Nguyen, Thao Nguyen, Fadillah Adamsyah Maani, Huy M. Le, Muhammad Umer Sheikh, Numan Saeed, Muhammad Haris Khan, Salman Khan arxiv

Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, satellite imagery, geospatial context, and simulator outputs. Weather and climate foundation models can forecast well, but do not reason interactively in language, while large language models (LLMs) reason in language but cannot operate directly on high-dimensional Earth-system data. As a result, real scientific workflows in Earth-science remain underserved. We introduce TerraBench, a benchmark for grounded Earth-science reasoning, built on TerraAgent, a ReAct-style executable framework that interleaves reasoning, tool calls, and observations to couple LLM planning with scientific tools for environmental retrieval, geospatial processing, simulation, and artifact-backed computation. TerraBench unifies analysis of Earth observation imagery, gridded data, GIS reasoning and simulation in a single executable interface, whereas prior benchmarks isolate these capabilities into narrow individual tasks. It is also the first in this space to pair process-level tool-use metrics with tolerance-aware numeric scoring. The benchmark comprises 403 extensive agentic tasks across three tracks (Fundamentals, Simulator-Grounded, and Document-Grounded Verification) and eight application domains with 24,500 verified execution steps. These results indicate that reliable Earth-science agents must go beyond tool access to coordinate heterogeneous workflows, parameterize tools precisely, and preserve artifact provenance.

📄 PDF Abstract BibTeX arXiv:2606.13148

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

2026-08-24 · Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang 외 arxiv

Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change est…

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

2026-02-19 · Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir, Hiyam Debary 외 arxiv

Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities to remote sensing remains challenging, as…

Multimodal Reasoning

AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems

2026-01-16 · Weiyi Wang, Xinchi Chen, Jingjing Gong, Xuanjing Huang 외 arxiv

Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent benchmarks largely focus on symbolic or we…

OpenEarth-Agent: From Tool Calling to Tool Creation for Open-Environment Earth Observation

2026-03-23 · Sijie Zhao, Feng Liu, Xueliang Zhang, Hao Chen 외 arxiv

Earth Observation (EO) is essential for perceiving dynamic land surface changes, yet deploying autonomous EO in open environments is hindered by the immense diversity of multi-source data and heterogeneous tasks. While r…

Earth Science Foundation Models: From Perception to Reasoning and Discovery

2026-05-09 · Xiangyu Zhao, Bo Liu, Yuehan Zhang, Zelin Song 외 arxiv

Large foundation models (FMs) are transforming Earth science by integrating heterogeneous multimodal data, such as multi-platform imagery, gridded reanalysis data, diverse geophysical and geochemical observations, and do…

Multimodal Reasoning