paper-with-me

홈 › Papers

Revisiting Pretraining Objectives for Tabular Deep Learning

2022-07-07 · Ivan Rubachev, Artem Alekberov, Yury Gorishniy, Artem Babenko

Recent deep learning models for tabular data currently compete with the traditional ML models based on decision trees (GBDT). Unlike GBDT, deep models can additionally benefit from pretraining, which is a workhorse of DL for vision and NLP. For tabular problems, several pretraining methods were proposed, but it is not entirely clear if pretraining provides consistent noticeable improvements and what method should be used, since the methods are often not compared to each other or comparison is limited to the simplest MLP architectures. In this work, we aim to identify the best practices to pretrain tabular DL models that can be universally applied to different datasets and architectures. Among our findings, we show that using the object target labels during the pretraining stage is beneficial for the downstream performance and advocate several target-aware pretraining objectives. Overall, our experiments demonstrate that properly performed pretraining significantly increases the performance of tabular DL models, which often leads to their superiority over GBDTs.

📄 PDF Abstract BibTeX arXiv:2207.03208

Code (2)

puhsu/tabular-dl-pretrain-objectives 공식 구현 pytorch
kalelpark/DeepLearning-for-Tabular-Data pytorch

Tasks

Deep Learning

Similar Papers 제목 키워드 기반

XTab: Cross-table Pretraining for Tabular Transformers

2023-05-10 · Bingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li 외

The success of self-supervised learning in computer vision and natural language processing has motivated pretraining methods on tabular data. However, most existing tabular self-supervised learning models fail to leverag…

AutoMLFederated LearningSelf-Supervised Learning

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

2025-11-20 · Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao 외 arxiv

Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, there is no consensus on whether audio-language models can build effective general-p…

Representation LearningContrastive Learning

Towards Cross-Table Masked Pretraining for Web Data Mining

2023-07-10 · Chao Ye, Guoshan Lu, Haobo Wang, Liyao Li 외

Tabular data pervades the landscape of the World Wide Web, playing a foundational role in the digital architecture that underpins online information. Given the recent influence of large-scale pretrained models like ChatG…

Contrastive Learning

Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data

2024-06-01 · Jintai Chen, Zhen Lin, Qiyuan Chen, Jimeng Sun

Tabular data from different tables exhibit significant diversity due to varied definitions and types of features, as well as complex inter-feature and feature-target relationships. Cross-dataset pretraining, which learns…

TabICLv2: A better, faster, scalable, and open tabular foundation model

2026-02-11 · Jingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le Morvan arxiv

Tabular foundation models, such as TabPFNv2 and TabICL, have recently dethroned gradient-boosted trees at the top of predictive benchmarks, demonstrating the value of in-context learning for tabular data. We introduce Ta…

Synthetic Data Generation