paper-with-me

Papers

Hybrid Data can Enhance the Utility of Synthetic Data for Training Anti-Money Laundering Models

2025-09-23 · Rachel Chung, Pratyush Nidhi Sharma, Mikko Siponen, Rohit Vadodaria, Luke Smith arxiv

Money laundering is a critical global issue for financial institutions. Automated Anti-money laundering (AML) models, like Graph Neural Networks (GNN), can be trained to identify illicit transactions in real time. A major issue for developing such models is the lack of access to training data due to privacy and confidentiality concerns. Synthetically generated data that mimics the statistical properties of real data but preserves privacy and confidentiality has been proposed as a solution. However, training AML models on purely synthetic datasets presents its own set of challenges. This article proposes the use of hybrid datasets to augment the utility of synthetic datasets by incorporating publicly available, easily accessible, and real-world features. These additions demonstrate that hybrid datasets not only preserve privacy but also improve model utility, offering a practical pathway for financial institutions to enhance AML systems.

📄 PDF Abstract BibTeX arXiv:2509.18499

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DSF-GAN: DownStream Feedback Generative Adversarial Network

2024-03-27 · Oriel Perets, Nadav Rappoport

Utility and privacy are two crucial measurements of the quality of synthetic tabular data. While significant advancements have been made in privacy measures, generating synthetic samples with high utility remains challen…

Generative Adversarial Network

SMOTE-DP: Improving Privacy-Utility Tradeoff with Synthetic Data

2025-06-02 · Yan Zhou, Bradley Malin, Murat Kantarcioglu

Privacy-preserving data publication, including synthetic data sharing, often experiences trade-offs between privacy and utility. Synthetic data is generally more effective than data anonymization in balancing this trade-…

Privacy PreservingSynthetic Data Generation

Hybrid Training Approaches for LLMs: Leveraging Real and Synthetic Data to Enhance Model Performance in Domain-Specific Applications

2024-10-11 · Alexey Zhezherau, Alexei Yanockin

This research explores a hybrid approach to fine-tuning large language models (LLMs) by integrating real-world and synthetic data to boost model performance, particularly in generating accurate and contextually relevant …

Diversity

Improve Fidelity and Utility of Synthetic Credit Card Transaction Time Series from Data-centric Perspective

2024-01-01 · Din-Yin Hsieh, Chi-Hua Wang, Guang Cheng

Exploring generative model training for synthetic tabular data, specifically in sequential contexts such as credit card transaction data, presents significant challenges. This paper addresses these challenges, focusing o…

Fraud DetectionTime Series

Structural MRI Synthesis for Alzheimer's Disease via Conditional Diffusion on Anatomical Masks

2026-06-16 · Muge Zhang, Muhammad Ali Khaliq, Jamal Alsakran, Byeong Kil Lee 외 arxiv

Recent advances in generative machine learning models have significantly improved medical imaging, offering promising solutions for data augmentation, privacy preservation, and improved model generalization. However, syn…

Data Augmentation