Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change
Climate-Eval is a comprehensive benchmark designed to evaluate natural language processing models across a broad range of tasks related to climate change. Climate-Eval aggregates existing datasets along with a newly developed news classification dataset, created specifically for this release. This results in a benchmark of 25 tasks based on 13 datasets, covering key aspects of climate discourse, including text classification, question answering, and information extraction. Our benchmark provides a standardized evaluation suite for systematically assessing the performance of large language models (LLMs) on these tasks. Additionally, we conduct an extensive evaluation of open-source LLMs (ranging from 2B to 70B parameters) in both zero-shot and few-shot settings, analyzing their strengths and limitations in the domain of climate change.
Code (0)
등록된 구현이 없습니다.
Tasks
News ClassificationQuestion Answeringtext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Towards Fine-grained Classification of Climate Change related Social Media Text
With climate change becoming a cause of concern worldwide, it becomes essential to gauge people’s reactions. This can help educate and spread awareness about it and help leaders improve decision-making. This work explore…
ClassificationDecision Makingnamed-entity-recognitionNamed Entity Recognition+5Adapting to climate change: Long-term impact of wind resource changes on China's power system resilience
Modern society's reliance on power systems is at risk from the escalating effects of wind-related climate change. Yet, failure to identify the intricate relationship between wind-related climate risks and power systems c…
Climate Change from Large Language Models
Climate change poses grave challenges, demanding widespread understanding and low-carbon lifestyle awareness. Large language models (LLMs) offer a powerful tool to address this crisis, yet comprehensive evaluations of th…
Prompt EngineeringClimaQA: An Automated Evaluation Framework for Climate Foundation Models
The use of foundation models in climate science has recently gained significant attention. However, a critical issue remains: the lack of a comprehensive evaluation framework capable of assessing the quality and scientif…
CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
Climate science demands automated workflows to transform comprehensive questions into data-driven statements across massive, heterogeneous datasets. However, generic LLM agents and static scripting pipelines lack climate…