paper-with-me

홈 › Papers

Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search

2021-06-25 · Junwen Yang, Yeye He, Surajit Chaudhuri

Recent work has made significant progress in helping users to automate single data preparation steps, such as string-transformations and table-manipulation operators (e.g., Join, GroupBy, Pivot, etc.). We in this work propose to automate multiple such steps end-to-end, by synthesizing complex data pipelines with both string transformations and table-manipulation operators. We propose a novel "by-target" paradigm that allows users to easily specify the desired pipeline, which is a significant departure from the traditional by-example paradigm. Using by-target, users would provide input tables (e.g., csv or json files), and point us to a "target table" (e.g., an existing database table or BI dashboard) to demonstrate how the output from the desired pipeline would schematically "look like". While the problem is seemingly underspecified, our unique insight is that implicit table constraints such as FDs and keys can be exploited to significantly constrain the space to make the problem tractable. We develop an Auto-Pipeline system that learns to synthesize pipelines using reinforcement learning and search. Experiments on large numbers of real pipelines crawled from GitHub suggest that Auto-Pipeline can successfully synthesize 60-70% of these complex pipelines with up to 10 steps.

📄 PDF Abstract BibTeX arXiv:2106.13861

Code (1)

https://gitlab.com/jwjwyoung/autopipeline-benchmarks 공식 구현

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

SapientML: Synthesizing Machine Learning Pipelines by Learning from Human-Written Solutions

2022-02-18 · Ripon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu 외

Automatic machine learning, or AutoML, holds the promise of truly democratizing the use of machine learning (ML), by substantially automating the work of data scientists. However, the huge combinatorial search space of c…

AutoMLBIG-bench Machine LearningProgram Synthesis

Visus: An Interactive System for Automatic Machine Learning Model Building and Curation

2019-07-05 · Aécio Santos, Sonia Castelo, Cristian Felix, Jorge Piazentin Ono 외

While the demand for machine learning (ML) applications is booming, there is a scarcity of data scientists capable of building such models. Automatic machine learning (AutoML) approaches have been proposed that help with…

AutoMLBIG-bench Machine Learning

AVATAR -- Machine Learning Pipeline Evaluation Using Surrogate Model

2020-01-30 · Tien-Dung Nguyen, Tomasz Maszczyk, Katarzyna Musial, Marc-Andre Zöller 외

The evaluation of machine learning (ML) pipelines is essential during automatic ML pipeline composition and optimisation. The previous methods such as Bayesian-based and genetic-based optimisation, which are implemented …

BIG-bench Machine Learningmodel

Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization

2026-06-25 · Janik Bischoff, Anne Meyer, Uta Mohring, Fabian Dunke 외 arxiv

Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picker routing. While integrated models capture interactions between these decisions, …

A Systematic Evaluation of Retrieval-Augmented Generation and Language Models for Space Operations

2026-05-23 · Ruben Belo, Marta Guimarães, Cláudia Soares arxiv

The rapid expansion of space activities has led to an unprecedented accumulation of technical documentation, operational guidelines, and scientific literature, creating challenges for timely decision-making in space oper…

Information Retrieval