paper-with-me

Papers

SAFE: A Novel Approach to AI Weather Evaluation through Stratified Assessments of Forecasts over Earth

2025-10-30 · Nick Masi, Randall Balestriero arxiv

The dominant paradigm in machine learning is to assess model performance based on average loss across all samples in some test set. This amounts to averaging performance geospatially across the Earth in weather and climate settings, failing to account for the non-uniform distribution of human development and geography. We introduce Stratified Assessments of Forecasts over Earth (SAFE), a package for elucidating the stratified performance of a set of predictions made over Earth. SAFE integrates various data domains to stratify by different attributes associated with geospatial gridpoints: territory (usually country), global subregion, income, and landcover (land or water). This allows us to examine the performance of models for each individual stratum of the different attributes (e.g., the accuracy in every individual country). To demonstrate its importance, we utilize SAFE to benchmark a zoo of state-of-the-art AI-based weather prediction models, finding that they all exhibit disparities in forecasting skill across every attribute. We use this to seed a benchmark of model forecast fairness through stratification at different lead times for various climatic variables. By moving beyond globally-averaged metrics, we for the first time ask: where do models perform best or worst, and which models are most fair? To support further work in this direction, the SAFE package is open source and available at https://github.com/N-Masi/safe

📄 PDF Abstract BibTeX arXiv:2510.26099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-Calibrated Automated Testing and Validation of Generative Language Models

2024-11-25 · Agus Sudjianto, Aijun Zhang, Srinivas Neppalli, Tarun Joshi 외

This paper introduces a comprehensive framework for the evaluation and validation of generative language models (GLMs), with a focus on Retrieval-Augmented Generation (RAG) systems deployed in high-stakes domains such as…

Conformal PredictionRAGRetrieval-augmented Generation

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

2025-11-19 · Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Francesco Giarrusso 외 arxiv

We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded…

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

2026-08-06 · Ro Encarnación, Tina Behzad, Emma Lurie, Danaé Metaxa arxiv

Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform…

RepIt: Steering Language Models with Concept-Specific Refusal Vectors

2025-09-16 · Vincent Siu, Nathan W. Henry, Nicholas Crispino, Yang Liu 외 arxiv

Current safety evaluations of language models rely on benchmark-based assessments that may miss localized vulnerabilities. We present RepIt, a simple and data-efficient framework for isolating concept-specific representa…

ClimaEmpact: Domain-Aligned Small Language Models and Datasets for Extreme Weather Analytics

2025-04-27 · Deeksha Varshney, Keane Ong, Rui Mao, Erik Cambria 외

Accurate assessments of extreme weather events are vital for research and policy, yet localized and granular data remain scarce in many parts of the world. This data gap limits our ability to analyze potential outcomes a…

ArticlesEmotion Recognition