paper-with-me

홈 › Papers

LLM-as-a-Discriminator: When Synthetic Tables Still Look Real

2026-06-01 · Manel Slokom, Malek Slokom, Thierno Kante arxiv

Privacy and data sharing are often in tension. Many organizations use synthetic data to reduce privacy risk and still share useful data. For tabular data, auditing privacy remains hard. In many cases, even humans cannot easily tell if a table is real or synthetic. In this paper, we propose a method based on LLM discrimination. We ask an LLM to classify each table sample as REAL or SYNTHETIC. We test two settings: C1 with table only, and C2 with table plus distributional metadata. We use LLaMA as an open model and Gemini as a reference model. In our experiments, we run three synthesis models, CTGAN, TVAE, and Gaussian Copula, on two public datasets, UCI Adult and ACS Census. We collect 451 valid trials. Our results show clear differences between models. On Adult, LLaMA reaches DRS=0% in reported cells, while Gemini reaches DRS=100% for CTGAN and TVAE. On Census, LLaMA predicts SYNTHETIC for most samples, while Gemini stays high in C1 but drops for CTGAN and TVAE in C2. We also compare with a classifier two-sample test (C2ST) and record linkage as distributional baselines, and with a human pilot of 2 annotators and 240 trials. Our results show that LLM discrimination is a practical privacy audit signal when model choice, per provider reporting, and data encoding are handled with care. For reproducibility, code and experiment scripts are available at https://github.com/SlokomManel/LLM-as-a-Discriminator.

📄 PDF Abstract BibTeX arXiv:2606.09865

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mind the Gap? A Distributional Comparison of Real and Synthetic Priors for Tabular Foundation Models

2026-05-07 · Alex O. Davies, Telmo de Menezes e Silva Filho, Nirav Ajmeri arxiv

Tabular foundation models are pre-trained on one of three classes of corpus: curated datasets drawn from benchmark repositories, tables harvested at scale from the web, or synthetic tables sampled from a parametric gener…

Learning Series-Parallel Lookup Tables for Efficient Image Super-Resolution

2022-07-26 · Cheng Ma, Jingyi Zhang, Jie zhou, Jiwen Lu

Lookup table (LUT) has shown its efficacy in low-level vision tasks due to the valuable characteristics of low computational cost and hardware independence. However, recent attempts to address the problem of single image…

Image Super-ResolutionSuper-Resolution

Learning View Priors for Single-view 3D Reconstruction

2018-11-26 · CVPR 2019 6 · Hiroharu Kato, Tatsuya Harada

There is some ambiguity in the 3D shape of an object when the number of observed views is small. Because of this ambiguity, although a 3D object reconstructor can be trained using a single view or a few views per object,…

3D ReconstructionObjectSingle-View 3D Reconstruction

LUT-Fuse: Towards Extremely Fast Infrared and Visible Image Fusion via Distillation to Learnable Look-Up Tables

2025-08-30 · Xunpeng Yi, Yibing Zhang, Xinyu Xiang, Qinglong Yan 외 arxiv

Current advanced research on infrared and visible image fusion primarily focuses on improving fusion performance, often neglecting the applicability on real-time fusion devices. In this paper, we propose a novel approach…

Cascade hash tables: a series of multilevel double hashing schemes with O(1) worst case lookup time

2015-06-25 · Li Shaohua

In this paper, the author proposes a series of multilevel double hashing schemes called cascade hash tables. They use several levels of hash tables. In each table, we use the common double hashing scheme. Higher level ha…