paper-with-me

홈 › Papers

Causal-Privacy Audit Workflow for Synthetic and Distilled Data in Dropout Support

2026-06-14 · Hanghang Zheng, Xiwei Zhuang, Zhong Wang, Hong Liu, Xiao Chen, Jingwen He, Xia Li arxiv

Synthetic and distilled student data are increasingly used to enable privacy-conscious learning analytics, yet their suitability for decision-facing institutional support remains uncertain. In dropout support, generated data must preserve not only predictive utility or distributional resemblance, but also the financial-status evidence used to guide advising, payment-plan assistance, and scholarship-related decisions. Method: This study introduces CaP-Eval, a decision-facing causal-privacy audit workflow for evaluating generated student data under a fixed estimand, timing-aware adjustment design, estimator set, and empirical privacy-governance screen. The workflow compares original, distilled, adversarial synthetic, statistical synthetic, and DPGNet privacy-oriented generated data on predictive utility, treatment-effect fidelity, robustness to alternative estimators, and local training-record proximity. Results: DPGNet and distilled data preserved the original financial-status treatment-effect structure more reliably than the adversarial and Gaussian Copula baselines. DPGNet preserved full direction and rank agreement across epsilon levels; epsilon = 10 produced the smallest non-original IPW and DML deviations, while epsilon = 1 and epsilon = 5 amplified several financial-status contrasts. Distilled data remained highly faithful but retained the strongest local training-record proximity signal. TabularGNet preserved qualitative directions with moderate attenuation, and Gaussian Copula compressed effect magnitudes. Conclusions: Predictive utility, privacy orientation, empirical disclosure signals, and causal fidelity diverged; generated student data require joint audits of direction, magnitude, overlap, and release-governance risk before decision use.

📄 PDF Abstract BibTeX arXiv:2606.15940

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents

2026-06-26 · Shijing Hu, Liang Liu, Zhu Meng, Zhicheng Zhao arxiv

Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments, and execute multi-step tasks. However, conventional function-calling…

Text Generation

Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data

2026-06-15 · Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard 외 arxiv

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. However, generating high-utility synthetic…

Synthetic Data Generation

Advancing the State-of-the-Art in Empirical Privacy Auditing

2026-06-09 · Nicole Mitchell, Galen Andrew, Arun Ganesh, Brendan McMahan 외 arxiv

Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples. Empirical privacy auditing (EPA) quantifies this risk by measuring realistic data leak…

parameter-efficient fine-tuning

Device-Native Autonomous Agents for Privacy-Preserving Negotiations

2026-01-01 · Joyjit Roy, Samaresh Kumar Singh arxiv

Automated negotiations in insurance and business-to-business (B2B) commerce encounter substantial challenges. Current systems force a trade-off between convenience and privacy by routing sensitive financial data through …

Quantitative Auditing of AI Fairness with Differentially Private Synthetic Data

2025-04-30 · Chih-Cheng Rex Yuan, Bow-Yaw Wang

Fairness auditing of AI systems can identify and quantify biases. However, traditional auditing using real-world data raises security and privacy concerns. It exposes auditors to security risks as they become custodians …

FairnessPrivacy Preserving