A Framework for Evaluating PM2.5 Forecasts from the Perspective of Individual Decision Making
Wildfire frequency is increasing as the climate changes, and the resulting air pollution poses health risks. Just as people routinely use weather forecasts to plan their activities around precipitation, reliable air quality forecasts could help individuals reduce their exposure to air pollution. In the present work, we evaluate several existing forecasts of fine particular matter (PM2.5) within the continental United States in the context of individual decision-making. Our comparison suggests there is meaningful room for improvement in air pollution forecasting, which might be realized by incorporating more data sources and using machine learning tools. To facilitate future machine learning development and benchmarking, we set up a framework to evaluate and compare air pollution forecasts for individual decision making. We introduce a new loss to capture decisions about when to use mitigation measures. We highlight the importance of visualizations when comparing forecasts. Finally, we provide code to download and compare archived forecast predictions.
Code (1)
Tasks
BenchmarkingDecision MakingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks
Standard weather forecast evaluations focus on the forecaster's perspective and on a statistical assessment comparing forecasts and observations. In practice, however, forecasts are used to make decisions, so it seems na…
Right Decisions from Wrong Predictions: A Mechanism Design Alternative to Individual Calibration
Decision makers often need to rely on imperfect probabilistic forecasts. While average performance metrics are typically available, it is difficult to assess the quality of individual forecasts and the corresponding util…
An Imbalance-Robust Evaluation Framework for Extreme Risk Forecasts
Evaluating rare-event forecasts is challenging because standard metrics collapse as event prevalence declines. Measures such as F1-score, AUPRC, MCC, and accuracy induce degenerate thresholds -- converging to zero or one…
Dynamic Asset Allocation with Asset-Specific Regime Forecasts
This article introduces a novel hybrid regime identification-forecasting framework designed to enhance multi-asset portfolio construction by integrating asset-specific regime forecasts. Unlike traditional approaches that…
Forecasts with Bayesian vector autoregressions under real time conditions
This paper investigates the sensitivity of forecast performance measures to taking a real time versus pseudo out-of-sample perspective. We use monthly vintages for the United States (US) and the Euro Area (EA) and estima…
Missing Values