paper-with-me

Papers

Empowering Tabular Data Preparation with Language Models: Why and How?

2025-08-03 · Mengshi Chen, Yuxiang Sun, Tengchao Li, Jianwei Wang, Kai Wang, Xuemin Lin, Ying Zhang, Wenjie Zhang arxiv

Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.

📄 PDF Abstract BibTeX arXiv:2508.01556

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lost in the Pipeline: How Well Do Large Language Models Handle Data Preparation?

2025-11-17 · Matteo Spreafico, Ludovica Tassini, Camilla Sancricca, Cinzia Cappiello arxiv

Large language models have recently demonstrated their exceptional capabilities in supporting and automating various tasks. Among the tasks worth exploring for testing large language model capabilities, we considered dat…

TablePilot: Recommending Human-Preferred Tabular Data Analysis with Large Language Models

2025-03-17 · Deyin Yi, Yihao Liu, Lang Cao, Mengyu Zhou 외

Tabular data analysis is crucial in many scenarios, yet efficiently identifying the most relevant data analysis queries and results for a new table remains a significant challenge. The complexity of tabular data, diverse…

SemPipes -- Optimizable Semantic Data Operators for Tabular Machine Learning Pipelines

2026-02-04 · Olga Ovcharenko, Matthias Boehm, Sebastian Schelter arxiv

Real-world machine learning on tabular data relies on complex data preparation pipelines for prediction, data integration, augmentation, and debugging. Designing these pipelines requires substantial domain expertise and …

Why Tabular Foundation Models Should Be a Research Priority

2024-05-02 · Boris van Breugel, Mihaela van der Schaar

Recent text and image foundation models are incredibly impressive, and these models are attracting an ever-increasing portion of research resources. In this position piece we aim to shift the ML research community's prio…

scientific discovery

Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey

2024-02-27 · Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang 외

Recent breakthroughs in large language modeling have facilitated rigorous exploration of their application in diverse tasks related to tabular data modeling, such as prediction, tabular data synthesis, question answering…

Language ModelingLanguage ModellingNavigateQuestion Answering+1