paper-with-me

홈 › Papers

SYNC: A Copula based Framework for Generating Synthetic Data from Aggregated Sources

2020-09-20 · Zheng Li, Yue Zhao, Jialin Fu

A synthetic dataset is a data object that is generated programmatically, and it may be valuable to creating a single dataset from multiple sources when direct collection is difficult or costly. Although it is a fundamental step for many data science tasks, an efficient and standard framework is absent. In this paper, we study a specific synthetic data generation task called downscaling, a procedure to infer high-resolution, harder-to-collect information (e.g., individual level records) from many low-resolution, easy-to-collect sources, and propose a multi-stage framework called SYNC (Synthetic Data Generation via Gaussian Copula). For given low-resolution datasets, the central idea of SYNC is to fit Gaussian copula models to each of the low-resolution datasets in order to correctly capture dependencies and marginal distributions, and then sample from the fitted models to obtain the desired high-resolution subsets. Predictive models are then used to merge sampled subsets into one, and finally, sampled datasets are scaled according to low-resolution marginal constraints. We make four key contributions in this work: 1) propose a novel framework for generating individual level data from aggregated data sources by combining state-of-the-art machine learning and statistical techniques, 2) perform simulation studies to validate SYNC's performance as a synthetic data generation algorithm, 3) demonstrate its value as a feature engineering tool, as well as an alternative to data collection in situations where gathering is difficult through two real-world datasets, 4) release an easy-to-use framework implementation for reproducibility and scalability at the production level that easily incorporates new data.

📄 PDF Abstract BibTeX arXiv:2009.09471

Code (1)

winstonll/SynC

Tasks

Feature EngineeringSynthetic Data Generation

Similar Papers 제목 키워드 기반

SynC: A Unified Framework for Generating Synthetic Population with Gaussian Copula

2019-04-16 · Colin Wan, Zheng Li, Alicia Guo, Yue Zhao

Synthetic population generation is the process of combining multiple socioeconomic and demographic datasets from different sources and/or granularity levels, and downscaling them to an individual level. Although it is a …

Feature Engineering

Copula-based synthetic data augmentation for machine-learning emulators

2020-12-16 · David Meyer, Thomas Nagler, Robin J. Hogan

Can we improve machine-learning (ML) emulators with synthetic data? If data are scarce or expensive to source and a physical model is available, statistically generated data may be useful for augmenting training sets che…

BIG-bench Machine LearningData AugmentationSynthetic Data Generation

Generation and Simulation of Synthetic Datasets with Copulas

2022-03-30 · Regis Houssou, Mihai-Cezar Augustin, Efstratios Rappos, Vivien Bonvin 외

This paper proposes a new method to generate synthetic data sets based on copula models. Our goal is to produce surrogate data resembling real data in terms of marginal and joint distributions. We present a complete and …

Differentially Private Non Parametric Copulas: Generating synthetic data with non parametric copulas under privacy guarantees

2024-09-27 · Pablo A. Osorio-Marulanda, John Esteban Castro Ramirez, Mikel Hernández Jiménez, Nicolas Moreno Reyes 외

Creation of synthetic data models has represented a significant advancement across diverse scientific fields, but this technology also brings important privacy considerations for users. This work focuses on enhancing a n…

Privacy PreservingSynthetic Data Generation

Copula estimation for nonsynchronous financial data

2019-04-23 · Arnab Chakrabarti, Rituparna Sen

Copula is a powerful tool to model multivariate data. We propose the modelling of intraday financial returns of multiple assets through copula. The problem originates due to the asynchronous nature of intraday financial …