paper-with-me

홈 › Papers

Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples

2023-07-27 · Peng Li, Yeye He, Cong Yan, Yue Wang, Surajit Chaudhuri

Relational tables, where each row corresponds to an entity and each column corresponds to an attribute, have been the standard for tables in relational databases. However, such a standard cannot be taken for granted when dealing with tables "in the wild". Our survey of real spreadsheet-tables and web-tables shows that over 30% of such tables do not conform to the relational standard, for which complex table-restructuring transformations are needed before these tables can be queried easily using SQL-based analytics tools. Unfortunately, the required transformations are non-trivial to program, which has become a substantial pain point for technical and non-technical users alike, as evidenced by large numbers of forum questions in places like StackOverflow and Excel/Power-BI/Tableau forums. We develop an Auto-Tables system that can automatically synthesize pipelines with multi-step transformations (in Python or other languages), to transform non-relational tables into standard relational forms for downstream analytics, obviating the need for users to manually program transformations. We compile an extensive benchmark for this new task, by collecting 244 real test cases from user spreadsheets and online forums. Our evaluation suggests that Auto-Tables can successfully synthesize transformations for over 70% of test cases at interactive speeds, without requiring any input from users, making this an effective tool for both technical and non-technical users to prepare data for analytics.

📄 PDF Abstract BibTeX arXiv:2307.14565

Code (1)

lipengcs/auto-tables-benchmark 공식 구현

Tasks

Attribute

Similar Papers 제목 키워드 기반

Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search

2021-06-25 · Junwen Yang, Yeye He, Surajit Chaudhuri

Recent work has made significant progress in helping users to automate single data preparation steps, such as string-transformations and table-manipulation operators (e.g., Join, GroupBy, Pivot, etc.). We in this work pr…

reinforcement-learningReinforcement Learning (RL)

ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models

2024-10-25 · Benjamin Newman, Yoonjoo Lee, Aakanksha Naik, Pao Siangliulue 외

When conducting literature reviews, scientists often create literature review tables - tables whose rows are publications and whose columns constitute a schema, a set of aspects used to compare and contrast the papers. C…

Synthesizing Realistic Data for Table Recognition

2024-04-17 · Qiyu Hou, Jun Wang, Meixuan Qiao, Lujun Tian

To overcome the limitations and challenges of current automatic table data annotation methods and random table data synthesis approaches, we propose a novel method for synthesizing annotation data specifically designed f…

Table annotationTable Recognition

TabulaX: Leveraging Large Language Models for Multi-Class Table Transformations

2024-11-26 · Arash Dargahi Nobari, Davood Rafiei

The integration of tabular data from diverse sources is often hindered by inconsistencies in formatting and representation, posing significant challenges for data analysts and personal digital assistants. Existing method…

Retrieve, Merge, Predict: Augmenting Tables with Data Lakes

2024-02-09 · Riccardo Cappuzzo, Aimee Coelho, Felix Lefebvre, Paolo Papotti 외

Machine-learning from a disparate set of tables, a data lake, requires assembling features by merging and aggregating tables. Data discovery can extend autoML to data tables by automating these steps. We present an in-de…

AutoMLBenchmarkingFeature Engineering