paper-with-me

Papers

TS-Arena -- A Live Forecast Pre-Registration Platform

2025-12-23 · Marcel Meyer, Sascha Kaltenpoth, Henrik Albers, Kevin Zalipski, Oliver Müller arxiv

Time Series Foundation Models (TSFMs) are transforming the field of forecasting. However, evaluating them on historical data is increasingly difficult due to the risks of train-test sample overlaps and temporal overlaps between correlated train and test time series. To address this, we introduce TS-Arena, a live forecasting platform that shifts evaluation from the known past to the unknown future. Building on the concept of continuous benchmarking, TS-Arena evaluates models on future data. Crucially, we introduce a strict forecasting pre-registration protocol: models must submit predictions before the ground-truth data physically exists. This makes test-set contamination impossible by design. The platform relies on a modular microservice architecture that harmonizes and structures data from different sources and orchestrates containerized model submissions. By enforcing a strict pre-registration protocol on live data streams, TS-Arena prevents information leakage offers a faster alternative to traditional static, infrequently repeated competitions (e.g. the M-Competitions). First empirical results derived from operating TS-Arena over one year of energy time series demonstrate that established TSFMs accumulate robust longitudinal scores over time, while the continuous nature of the benchmark simultaneously allows newcomers to demonstrate immediate competitiveness. TS-Arena provides the necessary infrastructure to assess the true generalization capabilities of modern forecasting models. The platform and corresponding code are available at https://ts-arena.live/.

📄 PDF Abstract BibTeX arXiv:2512.20761

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

2026-04-27 · Max Kleinebrahm, Jonathan Berrisch, Philipp Eiser, Wolf Fichtner 외 arxiv

Energy forecasting research faces a persistent comparability gap that makes it difficult to measure consistent progress over time. Reported accuracy gains are often not directly comparable because models are evaluated un…

Time Series Forecasting

Music Arena: Live Evaluation for Text-to-Music

2025-07-28 · Yonghyun Kim, Wayne Chi, Anastasios N. Angelopoulos, Wei-Lin Chiang 외 arxiv

We present Music Arena, an open platform for scalable human preference evaluation of text-to-music (TTM) models. Soliciting human preferences via listening studies is the gold standard for evaluation in TTM, but these st…

Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

2024-05-30 · Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 외

As LLMs continuously evolve, there is an urgent need for a reliable evaluation method that delivers trustworthy results promptly. Currently, static benchmarks suffer from inflexibility and unreliability, leading users to…

ChatbotFairness

VisionArena: 230K Real World User-VLM Conversations with Preference Labels

2024-12-11 · CVPR 2025 1 · Christopher Chou, Lisa Dunlap, Koki Mashita, Krishna Mandal 외

With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world con…

ChatbotSpatial Reasoning

Prediction Arena: Benchmarking AI Models on Real-World Prediction Markets

2026-03-28 · Jaden Zhang, Gardenia Liu, Oliver Johansson, Hileamlak Yitayew 외 arxiv

We introduce Prediction Arena, a benchmark for evaluating AI models' predictive accuracy and decision-making by enabling them to trade autonomously on live prediction markets with real capital. Unlike synthetic benchmark…

Computational Efficiency