paper-with-me

홈 › Papers

Understanding Context Sampling in TabPFN on Small Tabular Datasets

2026-07-29 · Mohammed Abdullah arxiv

TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the context size and which rows constitute the context. We study how these choices affect prediction stability, accuracy, and selection cost using repeated context sampling on 15 OpenML datasets. Specifically, we investigate (i) whether larger contexts reduce prediction variability across random draws, (ii) whether accuracy depends on preserving the training distribution or on feature-space coverage, and (iii) whether expensive selection methods such as K-Means and farthest-point sampling provide benefits over uniform random sampling. We find that larger contexts are both more accurate and substantially more stable, with AUC coefficient of variation decreasing from roughly 6 to 18% at k=16 to 1 to 4% at larger context sizes on datasets with room for improvement. Although accuracy correlates with distribution representativeness in random contexts, controlled experiments show that matching feature means alone can reduce accuracy by up to 0.5 AUC because it reduces context diversity. Mixed-effects analysis identifies diversity and coverage, rather than feature-mean matching, as the stronger predictor of accuracy (diversity beta=+0.23, p=3x10^-12; feature-mean shift beta=-0.01, p=0.71). K-Means and farthest-point sampling achieve similar accuracy to random selection while requiring two to three orders of magnitude more selection cost. These results show that random sampling succeeds because it provides feature-space coverage in expectation, not because it reproduces the underlying data distribution.

📄 PDF Abstract BibTeX arXiv:2607.26628

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Closer Look at TabPFN v2: Strength, Limitation, and Extension

2025-02-24 · Han-Jia Ye, Si-Yang Liu, Wei-Lun Chao

Tabular datasets are inherently heterogeneous, posing significant challenges for developing pre-trained foundation models. The recently introduced transformer-based Tabular Prior-data Fitted Network v2 (TabPFN v2) achiev…

In-Context Learning

TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention

2026-06-05 · Si-Yang Liu, Han-Jia Ye arxiv

Tabular foundation models, exemplified by TabPFN, perform prediction via in-context learning, inferring test labels directly from labeled training examples. They have demonstrated competitive performance, particularly on…

TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification Problems

2025-02-04 · Si-Yang Liu, Han-Jia Ye

TabPFN has emerged as a promising in-context learning model for tabular data, capable of directly predicting the labels of test samples given labeled training examples. It has demonstrated competitive performance, partic…

Computational EfficiencyIn-Context Learningtabular-classification

ConTextTab: A Semantics-Aware Tabular In-Context Learner

2025-06-12 · Marco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam Thelin

Tabular in-context learning (ICL) has recently achieved state-of-the-art (SOTA) performance on several tabular prediction tasks. Previously restricted to classification problems on small tables, recent advances such as T…

In-Context LearningWorld Knowledge

TabPFN-MT: A Natively Multitask In-Context Learner for Tabular Data

2026-05-16 · Cormac Cureton, Narges Armanfard arxiv

Prior-Data Fitted networks (PFNs) have been very successful in tabular contexts, handling prediction tasks in context. However, they are designed for single-task inference, meaning that predicting several target values w…

Computational Efficiency