Papers Tabular Data Generation
“Tabular Data Generation” 태그가 달린 논문 112편 · 필터 해제
Diffusion-Driven Synthetic Tabular Data Generation for Enhanced DoS/DDoS Attack Classification
Class imbalance refers to a situation where certain classes in a dataset have significantly fewer samples than oth- ers, leading to biased model performance. Class imbalance in network intrusion detection using Tabular D…
Network Intrusion DetectionTabular Data GenerationData AugmentationFraud DetectionExploring the Heterogeneity of Tabular Data: A Diversity-aware Data Generator via LLMs
Tabular data generation has become increasingly essential for enabling robust machine learning applications, which require large-scale, high-quality data. Existing solutions leverage generative models to learn original d…
Tabular Data GenerationWhen Tables Leak: Attacking String Memorization in LLM-Based Tabular Data Generation
Large Language Models (LLMs) have recently demonstrated remarkable performance in generating high-quality tabular synthetic data. In practice, two primary approaches have emerged for adapting LLMs to tabular data generat…
Synthetic Data GenerationTabular Data GenerationPrivacy Preserving Diffusion Models for Mixed-Type Tabular Data Generation
We introduce DP-FinDiff, a differentially private diffusion framework for synthesizing mixed-type tabular data. DP-FinDiff employs embedding-based representations for categorical features, reducing encoding overhead and …
Tabular Data GenerationInstruction Tuning of Large Language Models for Tabular Data Generation-in One Day
Tabular instruction tuning has emerged as a promising research direction for improving LLMs understanding of tabular data. However, the majority of existing works only consider question-answering and reasoning tasks over…
Tabular Data GenerationMalDataGen: A Modular Framework for Synthetic Tabular Data Generation in Malware Detection
High-quality data scarcity hinders malware detection, limiting ML performance. We introduce MalDataGen, an open-source modular framework for generating high-fidelity synthetic tabular data using modular deep learning mod…
Tabular Data GenerationMalware DetectionMembership Inference over Diffusion-models-based Synthetic Tabular Data
This study investigates the privacy risks associated with diffusion-based synthetic tabular data generation methods, focusing on their susceptibility to Membership Inference Attacks (MIAs). We examine two recent models, …
Synthetic Data GenerationTabular Data GenerationTowards Universal Debiasing for Language Models-based Tabular Data Generation
Large language models (LLMs) have achieved promising results in tabular data generation. However, inherent historical biases in tabular datasets often cause LLMs to exacerbate fairness issues, particularly when multiple …
Tabular Data GenerationLimited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes
Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient. Existing tabular generation approaches, such…
Tabular Data GenerationTAGAL: Tabular Data Generation using Agentic LLM Methods
The generation of data is a common approach to improve the performance of machine learning tasks, among which is the training of models for classification. In this paper, we present TAGAL, a collection of methods able to…
Tabular Data GenerationFairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples
Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require sp…
Synthetic Data GenerationTabular Data GenerationA Conditional GAN for Tabular Data Generation with Probabilistic Sampling of Latent Subspaces
The tabular form constitutes the standard way of representing data in relational database systems and spreadsheets. But, similarly to other forms, tabular data suffers from class imbalance, a problem that causes serious …
Tabular Data GenerationDependency-aware synthetic tabular data generation
Synthetic tabular data is increasingly used in privacy-sensitive domains such as health care, but existing generative models often fail to preserve inter-attribute relationships. In particular, functional dependencies (F…
Tabular Data GenerationDoubling Your Data in Minutes: Ultra-fast Tabular Data Generation via LLM-Induced Dependency Graphs
Tabular data is critical across diverse domains, yet high-quality datasets remain scarce due to privacy concerns and the cost of collection. Contemporary approaches adopt large language models (LLMs) for tabular augmenta…
Tabular Data GenerationDensity EstimationNot All Features Deserve Attention: Graph-Guided Dependency Learning for Tabular Data Generation with Language Models
Large Language Models (LLMs) have shown strong potential for tabular data generation by modeling textualized feature-value pairs. However, tabular data inherently exhibits sparse feature-level dependencies, where many fe…
Tabular Data GenerationGraph LearningRisk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation
Synthetic tabular data is essential for machine learning workflows, especially for expanding small or imbalanced datasets and enabling privacy-preserving data sharing. However, state-of-the-art generative models (GANs, V…
Tabular Data GenerationFASTGEN: Fast and Cost-Effective Synthetic Tabular Data Generation with LLMs
Synthetic data generation has emerged as an invaluable solution in scenarios where real-world data collection and usage are limited by cost and scarcity. Large language models (LLMs) have demonstrated remarkable capabili…
Synthetic Data GenerationTabular Data GenerationSynthetic Tabular Data Generation: A Comparative Survey for Modern Techniques
As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially for tabular datasets, which are central t…
Synthetic Data GenerationTabular Data GenerationGenerating Synthetic Relational Tabular Data via Structural Causal Models
Synthetic tabular data generation has received increasing attention in recent years, particularly with the emergence of foundation models for tabular data. The breakthrough success of TabPFN (Hollmann et al.,2025), which…
Tabular Data GenerationCausalDiffTab: Mixed-Type Causal-Aware Diffusion for Tabular Data Generation
Training data has been proven to be one of the most critical components in training generative AI. However, obtaining high-quality data remains challenging, with data privacy issues presenting a significant hurdle. To ad…
Tabular Data Generation