paper-with-me

홈 › Papers

Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources

2024-09-12 · Alisia Lupidi, Carlos Gemmell, Nicola Cancedda, Jane Dwivedi-Yu, Jason Weston, Jakob Foerster, Roberta Raileanu, Maria Lomeli

Large Language Models still struggle in challenging scenarios that leverage structured data, complex reasoning, or tool usage. In this paper, we propose Source2Synth: a new method that can be used for teaching LLMs new skills without relying on costly human annotations. Source2Synth takes as input a custom data source and produces synthetic data points with intermediate reasoning steps grounded in real-world sources. Source2Synth improves the dataset quality by discarding low-quality generations based on their answerability. We demonstrate the generality of this approach by applying it to two challenging domains: we test reasoning abilities in multi-hop question answering (MHQA), and tool usage in tabular question answering (TQA). Our method improves performance by 25.51% for TQA on WikiSQL and 22.57% for MHQA on HotPotQA compared to the fine-tuned baselines.

📄 PDF Abstract BibTeX arXiv:2409.08239

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-hop Question AnsweringQuestion AnsweringSynthetic Data Generation

Similar Papers 제목 키워드 기반

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

2026-06-19 · Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu, Pratinav Seth arxiv

Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality filtering as separate stages. This fragmen…

Synthetic Data Generation

How Good Are Synthetic Requirements ? Evaluating LLM-Generated Datasets for AI4RE

2025-06-26 · Abdelkarim El-Hajjami, Camille Salinesi

The shortage of publicly available, labeled requirements datasets remains a major barrier to advancing Artificial Intelligence for Requirements Engineering (AI4RE). While Large Language Models offer promising capabilitie…

Defect DetectionDiversitySynthetic Data Generation

FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples

2025-08-15 · Nitish Nagesh, Salar Shakibhamedan, Mahdi Bagheri, Ziyu Wang 외 arxiv

Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require sp…

Synthetic Data GenerationTabular Data Generation

On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey

2024-06-14 · Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao 외

Within the evolving landscape of deep learning, the dilemma of data quantity and quality has been a long-standing problem. The recent advent of Large Language Models (LLMs) offers a data-centric solution to alleviate the…

Synthetic Data Generation

Get away with less: Need of source side data curation to build parallel corpus for low resource Machine Translation

2026-01-13 · Saumitra Yadav, Manish Shrivastava arxiv

Data curation is a critical yet under-researched step in the machine translation training paradigm. To train translation systems, data acquisition relies primarily on human translations and digital parallel sources or, t…

Machine TranslationData Augmentation