paper-with-me

홈 › Papers

Balancing Cost and Effectiveness of Synthetic Data Generation Strategies for LLMs

2024-09-29 · Yung-Chieh Chan, George Pu, Apaar Shanker, Parth Suresh, Penn Jenks, John Heyer, Sam Denton

As large language models (LLMs) are applied to more use cases, creating high quality, task-specific datasets for fine-tuning becomes a bottleneck for model improvement. Using high quality human data has been the most common approach to unlock model performance, but is prohibitively expensive in many scenarios. Several alternative methods have also emerged, such as generating synthetic or hybrid data, but the effectiveness of these approaches remain unclear, especially in resource-constrained scenarios and tasks that are not easily verified. To investigate this, we group various synthetic data generation strategies into three representative categories -- Answer Augmentation, Question Rephrase and New Question -- and study the performance of student LLMs trained under various constraints, namely seed instruction set size and query budget. We demonstrate that these strategies are not equally effective across settings. Notably, the optimal data generation strategy depends strongly on the ratio between the available teacher query budget and the size of the seed instruction set. When this ratio is low, generating new answers to existing questions proves most effective, but as this ratio increases, generating new questions becomes optimal. Across all tasks, we find that choice of augmentation method and other design choices matter substantially more in low to mid data regimes than in high data regimes. We provide a practical framework for selecting the appropriate augmentation method across settings, taking into account additional factors such as the scalability of each method, the importance of verifying synthetic data, and the use of different LLMs for synthetic data generation.

📄 PDF Abstract BibTeX arXiv:2409.19759

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning

2026-07-07 · Sangwoo Lee, Sunghwan Park, Jaewoo Lee arxiv

Label skew in federated learning (FL) causes client drift and degrades global accuracy. Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost. We propo…

Federated LearningData Augmentation

On the Usefulness of Synthetic Tabular Data Generation

2023-06-27 · Dionysis Manousakas, Sergül Aydöre

Despite recent advances in synthetic data generation, the scientific community still lacks a unified consensus on its usefulness. It is commonly believed that synthetic data can be used for both data exchange and boostin…

Data AugmentationData SummarizationPrivacy PreservingSynthetic Data Generation+1

To SMOTE, or not to SMOTE?

2022-01-21 · Yotam Elor, Hadar Averbuch-Elor

Balancing the data before training a classifier is a popular technique to address the challenges of imbalanced binary classification in tabular data. Balancing is commonly achieved by duplication of minority samples or b…

Binary Classification

Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities

2025-10-13 · Urs Spiegelhalter, Jörg K. H. Franke, Frank Hutter arxiv

Adapting language models to new tasks through continued pretraining faces a fundamental trade-off: models must learn new capabilities while avoiding catastrophic forgetting of existing knowledge. While prior work has stu…

Synthetic Data GenerationGeneral Knowledge

Scalable Modular Synthetic Data Generation for Advancing Aerial Autonomy

2022-11-10 · Mehrnaz Sabet, Praveen Palanisamy, Sakshi Mishra

One major barrier to advancing aerial autonomy has been collecting large-scale aerial datasets for training machine learning models. Due to costly and time-consuming real-world data collection through deploying drones, t…

Data AugmentationSynthetic Data Generation