paper-with-me

홈 › Papers

Understanding the Surprising Generalization Properties of Tabular Foundation Models

2026-08-18 · Nour Shaheen, Junwei Ma, Alex Labach, Frank Hutter, Valentin Thomas, Anthony L. Caterini arxiv

Tabular Foundation Models (TFMs) increasingly rely on in-context learning, where a model receives labelled examples at inference time and predicts labels for new inputs without updating its weights. Existing TFMs are typically trained on either massive synthetic corpora or very large collections of real datasets. In contrast, we show that surprisingly strong transfer can emerge from self-supervised pre-training on just a single real table. In this setting, we also find that tables tend to be either broadly useful or broadly poor regardless of downstream prediction task, and that the strongest predictor of usefulness is the number of features rather than the number of instances. This leads to a task-centric interpretation of tabular pre-training: the number and the quality of tasks are essential for the pre-training of TFMs. We show that the same task-centric perspective can help corpus design at scale: fine-grained column-level pre-processing consistently improves downstream performance, while no improvements are observed when we filter or deduplicate at the dataset level. Finally, we offer a new perspective for how TFMs generalize: we believe that tabular in-context generalization is largely retrieval-based, and good models are those that learn to identify relevant examples in the provided context and aggregate them well. The mechanics of TFMs have been relatively understudied; our task-centric, retrieval-based perspective offers a new framework to guide future model and corpus design.

📄 PDF Abstract BibTeX arXiv:2608.17957

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generalization Can Emerge in Tabular Foundation Models From a Single Table

2025-11-12 · Junwei Ma, Nour Shaheen, Alex Labach, Amine Mhedhbi 외 arxiv

Deep tabular modelling increasingly relies on in-context learning where, during inference, a model receives a set of $(x,y)$ pairs as context and predicts labels for new inputs without weight updates. We challenge the pr…

Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization?

2025-12-09 · Jeongwhan Choi, Woosung Kang, Minseo Kim, Jongwoo Kim 외 arxiv

Foundation models pretrained on large data have demonstrated remarkable zero-shot generalization capabilities across domains. Building on the success of TabPFN for tabular data and its recent extension to time series, we…

Zero-shot GeneralizationFeature EngineeringNode Classification

Data Language Models: A New Foundation Model Class for Tabular Data

2026-05-07 · Eda Erol, Giuliano Pezzoli, Ozer Cem Kelahmet arxiv

Every major data modality now has a foundation model that understands it natively: text has language models, images have vision models, audio has audio models. Tabular data, the modality on which many consequential real-…

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

2026-05-19 · Xinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng 외 arxiv

Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are typically designed for \emph{non-strategic} settings where data dist…

Bringing Graphs to the Table: Zero-shot Node Classification via Tabular Foundation Models

2025-09-08 · Adrian Hayler, Xingyue Huang, İsmail İlkan Ceylan, Michael Bronstein 외 arxiv

Graph foundation models (GFMs) have recently emerged as a promising paradigm for achieving broad generalization across various graph data. However, existing GFMs are often trained on datasets that may not fully reflect r…

Time Series ForecastingNode ClassificationGraph Learning