Assessing Human Judgment Forecasts in the Rapid Spread of the Mpox Outbreak: Insights and Challenges for Pandemic Preparedness
In May 2022, mpox (formerly monkeypox) spread to non-endemic countries rapidly. Human judgment is a forecasting approach that has been sparsely evaluated during the beginning of an outbreak. We collected -- between May 19, 2022 and July 31, 2022 -- 1275 forecasts from 442 individuals of six questions about the mpox outbreak where ground truth data are now available. Individual human judgment forecasts and an equally weighted ensemble were evaluated, as well as compared to a random walk, autoregressive, and doubling time model. We found (1) individual human judgment forecasts underestimated outbreak size, (2) the ensemble forecast median moved closer to the ground truth over time but uncertainty around the median did not appreciably decrease, and (3) compared to computational models, for 2-8 week ahead forecasts, the human judgment ensemble outperformed all three models when using median absolute error and weighted interval score; for one week ahead forecasts a random walk outperformed human judgment. We propose two possible explanations: at the time a forecast was submitted, the mode was correlated with the most recent (and smaller) observation that would eventually determine ground truth. Several forecasts were solicited on a logarithmic scale which may have caused humans to generate forecasts with unintended, large uncertainty intervals. To aide in outbreak preparedness, platforms that solicit human judgment forecasts may wish to assess whether specifying a forecast on logarithmic scale matches an individual's intended forecast, support human judgment by finding cues that are typically used to build forecasts, and, to improve performance, tailor their platform to allow forecasters to assign zero probability to events.
Code (1)
Similar Papers 제목 키워드 기반
ClarityEthic: Explainable Moral Judgment Utilizing Contrastive Ethical Insights from Large Language Models
With the rise and widespread use of Large Language Models (LLMs), ensuring their safety is crucial to prevent harm to humans and promote ethical behaviors. However, directly assessing value valence (i.e., support or oppo…
Contrastive LearningInferring COVID-19 spreading rates and potential change points for case number forecasts
As COVID-19 is rapidly spreading across the globe, short-term modeling forecasts provide time-critical information for decisions on containment and mitigation strategies. A main challenge for short-term forecasts is the …
Bayesian InferenceIllusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-documented protocols -- details that are…
Text GenerationXForecast: Evaluating Natural Language Explanations for Time Series Forecasting
Time series forecasting aids decision-making, especially for stakeholders who rely on accurate predictions, making it very important to understand and explain these models to ensure informed decisions. Traditional explai…
Decision MakingTime SeriesTime Series ForecastingArgumentatively Coherent Judgmental Forecasting
Judgmental forecasting employs human opinions to make predictions about future events, rather than exclusively historical data as in quantitative forecasting. When these opinions form an argumentative structure around fo…