paper-with-me

홈 › Papers

Auto-Validate by-History: Auto-Program Data Quality Constraints to Validate Recurring Data Pipelines

2023-06-04 · Dezhan Tu, Yeye He, Weiwei Cui, Song Ge, Haidong Zhang, Han Shi, Dongmei Zhang, Surajit Chaudhuri

Data pipelines are widely employed in modern enterprises to power a variety of Machine-Learning (ML) and Business-Intelligence (BI) applications. Crucially, these pipelines are \emph{recurring} (e.g., daily or hourly) in production settings to keep data updated so that ML models can be re-trained regularly, and BI dashboards refreshed frequently. However, data quality (DQ) issues can often creep into recurring pipelines because of upstream schema and data drift over time. As modern enterprises operate thousands of recurring pipelines, today data engineers have to spend substantial efforts to \emph{manually} monitor and resolve DQ issues, as part of their DataOps and MLOps practices. Given the high human cost of managing large-scale pipeline operations, it is imperative that we can \emph{automate} as much as possible. In this work, we propose Auto-Validate-by-History (AVH) that can automatically detect DQ issues in recurring pipelines, leveraging rich statistics from historical executions. We formalize this as an optimization problem, and develop constant-factor approximation algorithms with provable precision guarantees. Extensive evaluations using 2000 production data pipelines at Microsoft demonstrate the effectiveness and efficiency of AVH.

📄 PDF Abstract BibTeX arXiv:2306.02421

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HAFixAgent: History-Aware Program Repair Agent

2025-11-02 · Yu Shi, Hao Li, Bram Adams, Ahmed E. Hassan arxiv

Automated program repair (APR) has recently shifted toward large language models and agent-based systems, yet most systems rely on local snapshot context, overlooking repository history. Prior work shows that repository …

Program Repair

CursorCore: Assist Programming through Aligning Anything

2024-10-09 · Hao Jiang, Qi Liu, Rui Li, Shengyu Ye 외

Large language models have been successfully applied to programming assistance tasks, such as code completion, code insertion, and instructional code editing. However, these applications remain insufficiently automated a…

Code Completion

Exploratory Experiments on Programming Autonomous Robots in Jadescript

2020-07-23 · Eleonora Iotti, Giuseppe Petrosino, Stefania Monica, Federico Bergenti

This paper describes exploratory experiments to validate the possibility of programming autonomous robots using an agent-oriented programming language. Proper perception of the environment, by means of various types of s…

TRACE: Early Detection of Chronic Kidney Disease Onset with Transformer-Enhanced Feature Embedding

2020-12-03 · Yu Wang, Ziqiao Guan, Wei Hou, Fusheng Wang

Chronic kidney disease (CKD) has a poor prognosis due to excessive risk factors and comorbidities associated with it. The early detection of CKD faces challenges of insufficient medical histories of positive patients and…

Disease PredictionPrognosis

AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management

2025-12-11 · Shizuo Tian, Hao Wen, Yuxuan Chen, Jiacheng Liu 외 arxiv

The rapid development of mobile GUI agents has stimulated growing research interest in long-horizon task automation. However, building agents for these tasks faces a critical bottleneck: the reliance on ever-expanding in…