paper-with-me

홈 › Papers

Rolling-Origin Validation Reverses Model Rankings in Multi-Step PM10 Forecasting: XGBoost, SARIMA, and Persistence

2026-03-19 · Federico Garcia Crespi, Eduardo Yubero Funes, Marina Alfosea Simon arxiv

(a) Many air quality forecasting studies report gains from machine learning, but evaluations often use static chronological splits and omit persistence baselines, so the operational added value under routine updating is unclear. (b) Using 2,350 daily PM10 observations from 2017 to 2024 at an urban background monitoring station in southern Europe, we compare XGBoost and SARIMA against persistence under a static split and a rolling-origin protocol with monthly updates. We report horizon-specific skill and the predictability horizon, defined as the maximum horizon with positive persistence-relative skill. Static evaluation suggests XGBoost performs well from one to seven days ahead, but rolling-origin evaluation reverses rankings: XGBoost is not consistently better than persistence at short and intermediate horizons, whereas SARIMA remains positively skilled across the full range. (c) For researchers, static splits can overstate operational usefulness and change rankings. For practitioners, rolling-origin, persistence-referenced skill profiles show which methods stay reliable at each lead time.

📄 PDF Abstract BibTeX arXiv:2603.20315

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Calibrated Dataset Condensation for Faster Hyperparameter Search

2024-05-27 · Mucong Ding, Yuancheng Xu, Tahseen Rabbani, Xiaoyu Liu 외

Dataset condensation can be used to reduce the computational cost of training multiple models on a large dataset by condensing the training dataset into a small synthetic set. State-of-the-art approaches rely on matching…

Dataset Condensation

Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

2026-08-26 · Zhexi Feng, Wuxi Chen, Bingrui Zhang arxiv

Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model…

Modeling and Reversing Brain Lesions Using Diffusion Models

2025-07-08 · Omar Zamzam, Haleh Akrami, Anand Joshi, Richard Leahy arxiv

Brain lesions are abnormalities or injuries in brain tissue that are often detectable using magnetic resonance imaging (MRI), which reveals structural changes in the affected areas. This broad definition of brain lesions…

Lesion Segmentation

Spectral Edge Dynamics of Training Trajectories: Signal--Noise Geometry Across Scales

2026-03-14 · Yongzhong Xu arxiv

Despite hundreds of millions of parameters, transformer training trajectories evolve within only a few coherent directions. We introduce Spectral Edge Dynamics (SED) to quantify this structure: a rolling-window SVD of pa…

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

2026-08-04 · Ine Gevers, Walter Daelemans arxiv

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability o…