paper-with-me

홈 › Papers

ITGPT: Generative Pretraining on Irregular Timeseries

2026-05-15 · Antoine Honoré, Ming Xiao arxiv

Timeseries regression models often struggle to leverage large volumes of labeled multimodal data, particularly when the data are irregularly sampled or contain missing values. This is common in domains like healthcare and predictive maintenance, where data are collected from unreliable sources, and labeling requires expert knowledge or costly equipments. Transformer-based large language models have proven effective on structured data such as text through self-supervised learning (SSL) and generative pretraining (GPT) frameworks. However, such models lack the flexibility to efficiently process irregularly sampled multimodal timeseries data. In this paper, we introduce ITGPT, an attention-based architecture designed for handling multimodal, irregularly sampled timeseries by allowing training with both SSL losses and GPT-like objectives. We evaluate its performance on a healthcare task with the TIHM dataset, and a predictive maintenance task with the CompX dataset. Our results demonstrate that ITGPT achieves state-of-the-art performance without requiring resampling, feature fusion or explicit data imputation. Furthermore, when labels are scarce, ITGPT effectively leverages unlabeled data through SSL and GPT training, outperforming the purely supervised approach. This represents an important step towards efficiently using large and unstructured timeseries datasets for practical inference tasks.

📄 PDF Abstract BibTeX arXiv:2605.16069

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

PAITS: Pretraining and Augmentation for Irregularly-Sampled Time Series

2023-08-25 · Nicasia Beebe-Wang, Sayna Ebrahimi, Jinsung Yoon, Sercan O. Arik 외

Real-world time series data that commonly reflect sequential human behavior are often uniquely irregularly sampled and sparse, with highly nonuniform sampling over time and entities. Yet, commonly-used pretraining and au…

Time Series

TransitGPT: A Generative AI-based framework for interacting with GTFS data using Large Language Models

2024-12-07 · Saipraneeth Devunuri, Lewis Lehe

This paper introduces a framework that leverages Large Language Models (LLMs) to answer natural language queries about General Transit Feed Specification (GTFS) data. The framework is implemented in a chatbot called Tran…

ChatbotNatural Language Queries

AuditGPT: Auditing Smart Contracts with ChatGPT

2024-04-05 · Shihao Xia, Shuai Shao, Mengting He, Tingting Yu 외

To govern smart contracts running on Ethereum, multiple Ethereum Request for Comment (ERC) standards have been developed, each containing a set of rules to guide the behaviors of smart contracts. Violating the ERC rules …

FADTI: Fourier and Attention Driven Diffusion for Multivariate Time Series Imputation

2025-12-17 · Runze Li, Hanchen Wang, Wenjie Zhang, Binghao Li 외 arxiv

Multivariate time series imputation is fundamental in applications such as healthcare, traffic forecasting, and biological modeling, where sensor failures and irregular sampling lead to pervasive missing values. However,…

Multivariate Time Series Imputation

SurF: A Generative Model for Multivariate Irregular Time Series Forecasting

2026-05-13 · Mohammad R. Rezaei, Tejas Balaji, Rahul G. Krishnan arxiv

Irregularly sampled multivariate event streams remain a difficult modality for generative modeling: tokenization-based approaches break down when inter-event intervals vary by orders of magnitude. We (i) propose \textbf{…

Time Series ForecastingPoint Processes