paper-with-me

Synthetic Data Evaluation

1개 벤치마크 · 논문 11편 · 이 태스크의 논문 보기 →

Benchmarks

Titanic

결과 7개

Most implemented

Papers

Reducing Instability in Synthetic Data Evaluation with a Super-Metric in MalDataGen

2025-11-20 · Anna Luiza Gomes da Silva, Diego Kreutz, Angelo Diniz, Rodrigo Mansilha 외 arxiv

Evaluating the quality of synthetic data remains a persistent challenge in the Android malware domain due to instability and the lack of standardization among existing metrics. This work integrates into MalDataGen a Supe…

Synthetic Data Evaluation

RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets

2025-10-07 · Jan Cegin, Branislav Pecher, Ivan Srba, Jakub Simko arxiv

LLMs are powerful generators of synthetic data, which are used for training smaller, specific models. This is especially valuable for low-resource languages, where human-labelled data is scarce but LLMs can still produce…

Synthetic Data Evaluation

Synth-MIA: A Testbed for Auditing Privacy Leakage in Tabular Data Synthesis

2025-09-22 · Joshua Ward, Xiaofeng Lin, Chi-Hua Wang, Guang Cheng arxiv

Tabular Generative Models are often argued to preserve privacy by creating synthetic datasets that resemble training data. However, auditing their empirical privacy remains challenging, as commonly used similarity metric…

Synthetic Data Evaluation

Struct-Bench: A Benchmark for Differentially Private Structured Text Generation

2025-09-12 · Shuaiqi Wang, Vikas Raunak, Arturs Backurs, Victor Reis 외 arxiv

Differentially private (DP) synthetic data generation is a promising technique for utilizing private datasets that otherwise cannot be exposed for model training or other analytics. While much research literature has foc…

Synthetic Data GenerationSynthetic Data EvaluationText Generation

An ELIXIR scoping review on domain-specific evaluation metrics for synthetic data in life sciences

2025-06-17 · Styliani-Christina Fragkouli, Somya Iqbal, Lisa Crossman, Barbara Gravel 외

Synthetic data has emerged as a powerful resource in life sciences, offering solutions for data scarcity, privacy protection and accessibility constraints. By creating artificial datasets that mirror the characteristics …

scientific discoverySynthetic Data Evaluation

What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models

2025-04-29 · Jan Kapar, Niklas Koenen, Martin Jullum

Evaluating synthetic tabular data is challenging, since they can differ from the real data in so many ways. There exist numerous metrics of synthetic data quality, ranging from statistical distances to predictive perform…

counterfactualFeature ImportanceSynthetic Data Evaluation

전체 11편 보기 →