paper-with-me

Papers

Scaling While Privacy Preserving: A Comprehensive Synthetic Tabular Data Generation and Evaluation in Learning Analytics

2024-01-12 · Qinyi Liu, Mohammad Khalil, Ronas Shakya, Jelena Jovanovic

Privacy poses a significant obstacle to the progress of learning analytics (LA), presenting challenges like inadequate anonymization and data misuse that current solutions struggle to address. Synthetic data emerges as a potential remedy, offering robust privacy protection. However, prior LA research on synthetic data lacks thorough evaluation, essential for assessing the delicate balance between privacy and data utility. Synthetic data must not only enhance privacy but also remain practical for data analytics. Moreover, diverse LA scenarios come with varying privacy and utility needs, making the selection of an appropriate synthetic data approach a pressing challenge. To address these gaps, we propose a comprehensive evaluation of synthetic data, which encompasses three dimensions of synthetic data quality, namely resemblance, utility, and privacy. We apply this evaluation to three distinct LA datasets, using three different synthetic data generation methods. Our results show that synthetic data can maintain similar utility (i.e., predictive performance) as real data, while preserving privacy. Furthermore, considering different privacy and data utility requirements in different LA scenarios, we make customized recommendations for synthetic data generation. This paper not only presents a comprehensive evaluation of synthetic data but also illustrates its potential in mitigating privacy concerns within the field of LA, thus contributing to a wider application of synthetic data in LA and promoting a better practice for open science.

📄 PDF Abstract BibTeX arXiv:2401.06883

Code (1)

ql909/mathematical_definitions 공식 구현

Tasks

Privacy PreservingSynthetic Data GenerationTabular Data Generation

Similar Papers 제목 키워드 기반

Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation

2026-04-08 · Qian Ma, Sarah Rajtmajer arxiv

Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy an…

Synthetic Data Generation

Privacy-Preserving Synthetic Review Generation with Diverse Writing Styles Using LLMs

2025-07-24 · Tevin Atwal, Chan Nam Tieu, Yefeng Yuan, Zhan Shi 외 arxiv

The increasing use of synthetic data generated by Large Language Models (LLMs) presents both opportunities and challenges in data-driven applications. While synthetic data provides a cost-effective, scalable alternative …

Deciphering the Interplay between Attack and Protection Complexity in Privacy-Preserving Federated Learning

2025-08-16 · Xiaojin Zhang, Mingcong Xu, Yiming Li, Wei Chen 외 arxiv

Federated learning (FL) offers a promising paradigm for collaborative model training while preserving data privacy. However, its susceptibility to gradient inversion attacks poses a significant challenge, necessitating r…

Federated Learning

Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data

2025-09-14 · Yunghwei Lai, Ziyue Wang, Weizhi Ma, Yang Liu arxiv

Synthetic data generation with Large Language Models (LLMs) has emerged as a promising solution in the medical domain to mitigate data scarcity and privacy constraints. However, existing approaches remain constrained by …

Synthetic Data Generation

Beyond Real Faces: Synthetic Datasets Can Achieve Reliable Recognition Performance without Privacy Compromise

2025-10-20 · Paweł Borsukiewicz, Fadi Boutros, Iyiola E. Olatunji, Charles Beumier 외 arxiv

The deployment of facial recognition systems has created an ethical dilemma: achieving high accuracy requires massive datasets of real faces collected without consent, leading to dataset retractions and potential legal l…