paper-with-me

홈 › Papers

Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes

2023-12-19 · Nabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der Schaar

Machine Learning (ML) in low-data settings remains an underappreciated yet crucial problem. Hence, data augmentation methods to increase the sample size of datasets needed for ML are key to unlocking the transformative potential of ML in data-deprived regions and domains. Unfortunately, the limited training set constrains traditional tabular synthetic data generators in their ability to generate a large and diverse augmented dataset needed for ML tasks. To address this challenge, we introduce CLLM, which leverages the prior knowledge of Large Language Models (LLMs) for data augmentation in the low-data regime. However, not all the data generated by LLMs will improve downstream utility, as for any generative model. Consequently, we introduce a principled curation mechanism, leveraging learning dynamics, coupled with confidence and uncertainty metrics, to obtain a high-quality dataset. Empirically, on multiple real-world datasets, we demonstrate the superior performance of CLLM in the low-data regime compared to conventional generators. Additionally, we provide insights into the LLM generation and curation mechanism, shedding light on the features that enable them to output high-quality augmented datasets.

📄 PDF Abstract BibTeX arXiv:2312.12112

Code (2)

seedatnabeel/cllm 공식 구현 pytorch
vanderschaarlab/cllm 공식 구현 pytorch

Tasks

Data Augmentation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

2025-07-02 · Chris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He 외 arxiv

Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced h…

Reinforcement Learning

Engineering Regression Without Real-Data Training: Domain Adaptation for Tabular Foundation Models Using Multi-Dataset Embeddings

2026-03-05 · Lyle Regenwetter, Rosen Yu, Cyril Picard, Faez Ahmed arxiv

Predictive modeling in engineering applications has long been dominated by bespoke models and small, siloed tabular datasets, limiting the applicability of large-scale learning approaches. Despite recent progress in tabu…

Domain Adaptation

Large Language Models as Automated Aligners for benchmarking Vision-Language Models

2023-11-24 · Yuanfeng Ji, Chongjian Ge, Weikai Kong, Enze Xie 외

With the advancements in Large Language Models (LLMs), Vision-Language Models (VLMs) have reached a new level of sophistication, showing notable competence in executing intricate cognition and reasoning tasks. However, e…

BenchmarkingWorld Knowledge

Curation Leaks: Membership Inference Attacks against Data Curation for Machine Learning

2026-02-28 · Dariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska Boenisch arxiv

In machine learning, curation is used to select the most valuable data for improving both model accuracy and computational efficiency. Recently, curation has also been explored as a solution for private machine learning:…

Computational Efficiency

Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks

2024-10-03 · Xiaoqun Liu, Jiacheng Liang, Luoxi Tang, Muchao Ye 외

Large language models (LLMs) are widely adapted for downstream applications through fine-tuning, a process named customization. However, recent studies have identified a vulnerability during this process, where malicious…