paper-with-me

홈 › Papers

CLINB: A Climate Intelligence Benchmark for Foundational Models

2025-10-29 · Michelle Chen Huebscher, Katharine Mach, Aleksandar Stanić, Markus Leippold, Ben Gaiarin, Zeke Hausfather, Elisa Rawat, Erich Fischer, Massimiliano Ciaramita, Joeri Rogelj, Christian Buck, Lierni Sestorain Saralegui, Reto Knutti arxiv

Evaluating how Large Language Models (LLMs) handle complex, specialized knowledge remains a critical challenge. We address this through the lens of climate change by introducing CLINB, a benchmark that assesses models on open-ended, grounded, multimodal question answering tasks with clear requirements for knowledge quality and evidential support. CLINB relies on a dataset of real users' questions and evaluation rubrics curated by leading climate scientists. We implement and validate a model-based evaluation process and evaluate several frontier models. Our findings reveal a critical dichotomy. Frontier models demonstrate remarkable knowledge synthesis capabilities, often exhibiting PhD-level understanding and presentation quality. They outperform "hybrid" answers curated by domain experts assisted by weaker models. However, this performance is countered by failures in grounding. The quality of evidence varies, with substantial hallucination rates for references and images. We argue that bridging this gap between knowledge synthesis and verifiable attribution is essential for the deployment of AI in scientific workflows and that reliable, interpretable benchmarks like CLINB are needed to progress towards building trustworthy AI systems.

📄 PDF Abstract BibTeX arXiv:2511.11597

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases

2025-05-30 · Yuchong Li, Xiaojun Zeng, Chihua Fang, Jian Yang 외

Hepato-pancreato-biliary (HPB) disorders represent a global public health challenge due to their high morbidity and mortality. Although large language models (LLMs) have shown promising performance in general medical que…

Medical Question AnsweringMultiple-choiceQuestion Answering

Randomized Algorithms for Scientific Computing (RASC)

2021-04-19 · Aydin Buluc, Tamara G. Kolda, Stefan M. Wild, Mihai Anitescu 외

Randomized algorithms have propelled advances in artificial intelligence and represent a foundational research area in advancing AI for Science. Future advancements in DOE Office of Science priority areas such as climate…

WxC-Bench: A Novel Dataset for Weather and Climate Downstream Tasks

2024-12-03 · Rajat Shinde, Christopher E. Phillips, Kumar Ankur, Aman Gupta 외

High-quality machine learning (ML)-ready datasets play a foundational role in developing new artificial intelligence (AI) models or fine-tuning existing models for scientific applications such as weather and climate anal…

ClimateBench-M: A Multi-Modal Climate Data Benchmark with a Simple Generative Method

2025-04-10 · Dongqi Fu, Yada Zhu, Zhining Liu, Lecheng Zheng 외

Climate science studies the structure and dynamics of Earth's climate system and seeks to understand how climate changes over time, where the data is usually stored in the format of time series, recording the climate fea…

Time SeriesWeather Forecasting

From General to Specialized: The Need for Foundational Models in Agriculture

2025-07-07 · Vishal Nedungadi, Xingguo Xiong, Aike Potze, Ron Van Bree 외 arxiv

Food security remains a global concern as population grows and climate change intensifies, demanding innovative solutions for sustainable agricultural productivity. Recent advances in foundation models have demonstrated …