paper-with-me

홈 › Papers

Quality Matters: Evaluating Synthetic Data for Tool-Using LLMs

2024-09-24 · Shadi Iskander, Nachshon Cohen, Zohar Karnin, Ori Shapira, Sofia Tolmach

Training large language models (LLMs) for external tool usage is a rapidly expanding field, with recent research focusing on generating synthetic data to address the shortage of available data. However, the absence of systematic data quality checks poses complications for properly training and testing models. To that end, we propose two approaches for assessing the reliability of data for training LLMs to use external tools. The first approach uses intuitive, human-defined correctness criteria. The second approach uses a model-driven assessment with in-context evaluation. We conduct a thorough evaluation of data quality on two popular benchmarks, followed by an extrinsic evaluation that showcases the impact of data quality on model performance. Our results demonstrate that models trained on high-quality data outperform those trained on unvalidated data, even when trained with a smaller quantity of data. These findings empirically support the significance of assessing and ensuring the reliability of training data for tool-using LLMs.

📄 PDF Abstract BibTeX arXiv:2409.16341

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

2026-05-18 · Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran 외 arxiv

Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information. Generating and evaluating synthetic data across privacy, ut…

Decision Making

Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents

2026-04-04 · Ádám Kovács arxiv

Coding agents repeatedly consume long tool observations even though only a small fraction of each observation matters for the next step. We study task-conditioned tool-output pruning: given a focused query and one tool o…

Quality Measures in Biometric Systems

2021-11-17 · Fernando Alonso-Fernandez, Julian Fierrez, Javier Ortega-Garcia

Biometric technology has been increasingly deployed in the past decade, offering greater security and convenience than traditional methods of personal recognition. Although biometric signals' quality heavily affects a bi…

An Automated Length-Aware Quality Metric for Summarization

2025-07-10 · Andrew D. Foland arxiv

This paper proposes NOrmed Index of Retention (NOIR), a quantitative objective metric for evaluating summarization quality of arbitrary texts that relies on both the retention of semantic meaning and the summary length c…

Semantic Similarity

SynQP: A Framework and Metrics for Evaluating the Quality and Privacy Risk of Synthetic Data

2026-01-17 · Bing Hu, Yixin Li, Asma Bahamyirou, Helen Chen arxiv

The use of synthetic data in health applications raises privacy concerns, yet the lack of open frameworks for privacy evaluations has slowed its adoption. A major challenge is the absence of accessible benchmark datasets…

Synthetic Data Generation