paper-with-me

Papers

Downstream Task-Oriented Generative Model Selections on Synthetic Data Training for Fraud Detection Models

2024-01-01 · Yinan Cheng, Chi-Hua Wang, Vamsi K. Potluru, Tucker Balch, Guang Cheng

Devising procedures for downstream task-oriented generative model selections is an unresolved problem of practical importance. Existing studies focused on the utility of a single family of generative models. They provided limited insights on how synthetic data practitioners select the best family generative models for synthetic training tasks given a specific combination of machine learning model class and performance metric. In this paper, we approach the downstream task-oriented generative model selections problem in the case of training fraud detection models and investigate the best practice given different combinations of model interpretability and model performance constraints. Our investigation supports that, while both Neural Network(NN)-based and Bayesian Network(BN)-based generative models are both good to complete synthetic training task under loose model interpretability constrain, the BN-based generative models is better than NN-based when synthetic training fraud detection model under strict model interpretability constrain. Our results provides practical guidance for machine learning practitioner who is interested in replacing their training dataset from real to synthetic, and shed lights on more general downstream task-oriented generative model selection problems.

📄 PDF Abstract BibTeX arXiv:2401.00974

Code (0)

등록된 구현이 없습니다.

Tasks

Fraud DetectionModel Selection

Similar Papers 제목 키워드 기반

Variational Hierarchical Dialog Autoencoder for Dialog State Tracking Data Augmentation

2020-01-23 · EMNLP 2020 11 · Kang Min Yoo, Hanbit Lee, Franck Dernoncourt, Trung Bui 외

Recent works have shown that generative data augmentation, where synthetic samples generated from deep generative models complement the training dataset, benefit NLP tasks. In this work, we extend this approach to the ta…

Data Augmentationdialog state trackingDialogue State TrackingResponse Generation+2

Task-Oriented Data Synthesis and Control-Rectify Sampling for Remote Sensing Semantic Segmentation

2025-12-18 · Yunkai Yang, Yudong Zhang, Kunquan Zhang, Jinxiao Zhang 외 arxiv

With the rapid progress of controllable generation, training data synthesis has become a promising way to expand labeled datasets and alleviate manual annotation in remote sensing (RS). However, the complexity of semanti…

Semantic Segmentation

Synthetic data, real errors: how (not) to publish and use synthetic data

2023-05-16 · Boris van Breugel, Zhaozhi Qian, Mihaela van der Schaar

Generating synthetic data through generative models is gaining interest in the ML community and beyond, promising a future where datasets can be tailored to individual needs. Unfortunately, synthetic data is usually not …

Uncertainty Quantification

Task Oriented In-Domain Data Augmentation

2024-06-24 · Xiao Liang, Xinyu Hu, Simiao Zuo, Yeyun Gong 외

Large Language Models (LLMs) have shown superior performance in various applications and fields. To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on …

Data AugmentationMath

LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection

2026-05-11 · Abhishek Moturu, Anna Goldenberg, Babak Taati arxiv

Synthetic data is useful only when the added samples fill missing parts of the training distribution that matter for the downstream task. We introduce LiBaGS, a lightweight, generator-agnostic method for targeted synthet…