paper-with-me

홈 › Papers

CARTE: Pretraining and Transfer for Tabular Learning

2024-02-26 · Myung Jun Kim, Léo Grinsztajn, Gaël Varoquaux

Pretrained deep-learning models are the go-to solution for images or text. However, for tabular data the standard is still to train tree-based models. Indeed, transfer learning on tables hits the challenge of data integration: finding correspondences, correspondences in the entries (entity matching) where different words may denote the same entity, correspondences across columns (schema matching), which may come in different orders, names... We propose a neural architecture that does not need such correspondences. As a result, we can pretrain it on background data that has not been matched. The architecture -- CARTE for Context Aware Representation of Table Entries -- uses a graph representation of tabular (or relational) data to process tables with different columns, string embedding of entries and columns names to model an open vocabulary, and a graph-attentional network to contextualize entries with column names and neighboring entries. An extensive benchmark shows that CARTE facilitates learning, outperforming a solid set of baselines including the best tree-based models. CARTE also enables joint learning across tables with unmatched columns, enhancing a small table with bigger ones. CARTE opens the door to large pretrained models for tabular data.

📄 PDF Abstract BibTeX arXiv:2402.16785

Code (1)

soda-inria/carte 공식 구현 pytorch

Tasks

Data IntegrationTransfer Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular Barriers

2025-12-16 · Yibing Fu, Yunpeng Zhao, Zhitao Zeng, Cheng Chen 외 arxiv

Multi-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining thes…

Self-Supervised LearningRepresentation Learning

Fine-Tuning the Retrieval Mechanism for Tabular Deep Learning

2023-11-13 · Felix den Breejen, Sangmin Bae, Stephen Cha, Tae-Young Kim 외

While interests in tabular deep learning has significantly grown, conventional tree-based models still outperform deep learning methods. To narrow this performance gap, we explore the innovative retrieval mechanism, a me…

Deep LearningRetrievalTransfer Learning

TransTab: Learning Transferable Tabular Transformers Across Tables

2022-05-19 · Zifeng Wang, Jimeng Sun

Tabular data (or tables) are the most widely used data format in machine learning (ML). However, ML models often assume the table structure keeps fixed in training and testing. Before ML modeling, heavy data cleaning is …

Incremental LearningTransfer Learning

UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science

2023-07-18 · Yazheng Yang, Yuqi Wang, Guang Liu, Ledell Wu 외

Recent advancements in NLP have witnessed the groundbreaking impact of pretrained models, yielding impressive outcomes across various tasks. This study seeks to extend the power of pretraining methodologies to facilitati…

Revisiting Pretraining Objectives for Tabular Deep Learning

2022-07-07 · Ivan Rubachev, Artem Alekberov, Yury Gorishniy, Artem Babenko

Recent deep learning models for tabular data currently compete with the traditional ML models based on decision trees (GBDT). Unlike GBDT, deep models can additionally benefit from pretraining, which is a workhorse of DL…

Deep Learning