paper-with-me

Papers

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

2025-09-27 · Wanjin Feng, Yuan Yuan, Jingtao Ding, Yong Li arxiv

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, this approach suffers from a fundamental flaw: standard evaluation metrics conflate a model's performance with the data's intrinsic unpredictability. To address this pressing challenge, we introduce a novel, predictability-aligned diagnostic framework grounded in spectral coherence. Our framework makes two primary contributions: the Spectral Coherence Predictability (SCP), a computationally efficient ($O(N\log N)$) and task-aligned score that quantifies the inherent difficulty of a given forecasting instance, and the Linear Utilization Ratio (LUR), a frequency-resolved diagnostic tool that precisely measures how effectively a model exploits the linearly predictable information within the data. We validate our framework's effectiveness and leverage it to reveal two core insights. First, we provide the first systematic evidence of "predictability drift", demonstrating that a task's forecasting difficulty varies sharply over time. Second, our evaluation reveals a key architectural trade-off: complex models are superior for low-predictability data, whereas linear models are highly effective on more predictable tasks. We advocate for a paradigm shift, moving beyond simplistic aggregate scores toward a more insightful, predictability-aware evaluation that fosters fairer model comparisons and a deeper understanding of model behavior.

📄 PDF Abstract BibTeX arXiv:2509.23074

Code (0)

등록된 구현이 없습니다.

Tasks

Time Series Forecasting

Similar Papers 제목 키워드 기반

Forecast Collapse in Time-Series Foundation Models

2026-08-14 · Shu Wan, Miles Ma, Hank Zhu, Guangqi Liu 외 arxiv

When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast co…

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

2026-02-12 · Zhongzheng Qiao, Sheng Pan, Anni Wang, Viktoriya Zhukova 외 arxiv

Time series foundation models (TSFMs) are revolutionizing the forecasting landscape from specific dataset modeling to generalizable task evaluation. However, we contend that existing benchmarks exhibit common limitations…

Time Series Forecasting

Rolling-Origin Validation Reverses Model Rankings in Multi-Step PM10 Forecasting: XGBoost, SARIMA, and Persistence

2026-03-19 · Federico Garcia Crespi, Eduardo Yubero Funes, Marina Alfosea Simon arxiv

(a) Many air quality forecasting studies report gains from machine learning, but evaluations often use static chronological splits and omit persistence baselines, so the operational added value under routine updating is …

CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction

2026-08-17 · Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li 외 arxiv

Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clinical variables, whole-slide images, genomic profiles, or multimodal inputs, while…

Quantifying the Effects of Word Length, Frequency, and Predictability on Dyslexia

2025-10-28 · Hugo Rydel-Johnston, Alex Kafkas arxiv

We ask where, and under what conditions, dyslexic reading costs arise in a large-scale naturalistic reading dataset. Using eye-tracking aligned to word-level features (word length, frequency, and predictability), we mode…