paper-with-me

홈 › Papers

Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts

2024-09-24 · Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, Ming Jin

Deep learning for time series forecasting has seen significant advancements over the past decades. However, despite the success of large-scale pre-training in language and vision domains, pre-trained time series models remain limited in scale and operate at a high cost, hindering the development of larger capable forecasting models in real-world applications. In response, we introduce Time-MoE, a scalable and unified architecture designed to pre-train larger, more capable forecasting foundation models while reducing inference costs. By leveraging a sparse mixture-of-experts (MoE) design, Time-MoE enhances computational efficiency by activating only a subset of networks for each prediction, reducing computational load while maintaining high model capacity. This allows Time-MoE to scale effectively without a corresponding increase in inference costs. Time-MoE comprises a family of decoder-only transformer models that operate in an auto-regressive manner and support flexible forecasting horizons with varying input context lengths. We pre-trained these models on our newly introduced large-scale data Time-300B, which spans over 9 domains and encompassing over 300 billion time points. For the first time, we scaled a time series foundation model up to 2.4 billion parameters, achieving significantly improved forecasting precision. Our results validate the applicability of scaling laws for training tokens and model size in the context of time series forecasting. Compared to dense models with the same number of activated parameters or equivalent computation budgets, our models consistently outperform them by large margin. These advancements position Time-MoE as a state-of-the-art solution for tackling real-world time series forecasting challenges with superior capability, efficiency, and flexibility.

📄 PDF Abstract BibTeX arXiv:2409.16040

Code (1)

time-moe/time-moe 공식 구현 pytorch

Tasks

Computational EfficiencyMixture-of-ExpertsTime SeriesTime Series Forecasting

Similar Papers 제목 키워드 기반

Empowering Time Series Analysis with Large-Scale Multimodal Pretraining

2026-02-05 · Peng Chen, Siyuan Wang, Shiyan Hu, Xingjian Wu 외 arxiv

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural…

Time Series ForecastingTime Series AnalysisAnomaly Detection

QuitoBench: A High-Quality Open Time Series Forecasting Benchmark

2026-03-27 · Siqiao Xue, Zhaoyang Zhu, Wei Zhang, Rongyao Cai 외 arxiv

Time series forecasting is critical across finance, healthcare, and cloud computing, yet progress is constrained by a fundamental bottleneck: the scarcity of large-scale, high-quality benchmarks. To address this gap, we …

Time Series Forecasting

Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling

2026-03-05 · Yong Liu, Xingjian Su, Shiyu Wang, Haoran Zhang 외 arxiv

We introduce Timer-S1, a strong Mixture-of-Experts (MoE) time series foundation model with 8.3B total parameters, 0.75B activated parameters for each token, and a context length of 11.5K. To overcome the scalability bott…

Data Augmentation

MIRA: Medical Time Series Foundation Model for Real-World Health Data

2025-06-09 · Hao Li, Bowen Deng, Chang Xu, Zhiyuan Feng 외

A unified foundation model for medical time series -- pretrained on open access and ethics board-approved medical corpora -- offers the potential to reduce annotation burdens, minimize model customization, and enable rob…

EthicsMissing ValuesMixture-of-ExpertsTime Series+1

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

2026-07-07 · Qian Sun, Yong-Ming Tian, Jia-Wei Huang, Cheng Feng 외 arxiv

Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero-shot generalization. Modern multivariate TSFMs are predominantly pretrained on m…

Zero-shot Generalization