paper-with-me

홈 › Papers

Autodata: An agentic data scientist to create high quality synthetic data

2026-06-24 · Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston arxiv

We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning with mathematical objects, where we obtain improved results compared to classical synthetic dataset creation methods. Further, meta-optimizing the data scientist agent itself delivers an even larger performance uplift. Agentic data creation provides a way to convert increased inference compute into higher quality model training. Overall, we believe this direction has the potential to change the way we build AI data.

📄 PDF Abstract BibTeX arXiv:2606.25996

Code (0)

등록된 구현이 없습니다.

Tasks

Legal Reasoning

Similar Papers 제목 키워드 기반

AutoData: A Multi-Agent System for Open Web Data Collection

2025-05-21 · Tianyi Ma, Yiyue Qian, Zheyuan Zhang, Zehong Wang 외

The exponential growth of data-driven systems and AI technologies has intensified the demand for high-quality web-sourced datasets. While existing datasets have proven valuable, conventional web data collection approache…

Large Language Model

Democratizing AI scientists using ToolUniverse

2025-09-27 · Shanghua Gao, Richard Zhu, Pengwei Sui, Zhenglun Kong 외 arxiv

AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build because they are bespoke, tied to rigid workflows, and lack shared environments …

HeurekaBench: A Benchmarking Framework for AI Co-scientist

2026-01-04 · Siba Smarak Panigrahi, Jovana Videnović, Maria Brbić arxiv

LLM-based reasoning models have enabled the development of agentic systems that act as co-scientists, assisting in multi-step scientific analysis. However, evaluating these systems is challenging, as it requires realisti…

AI, Humans, and Data Science: Optimizing Roles Across Workflows and the Workforce

2025-07-15 · Richard Timpone, Yongwei Yang arxiv

AI is transforming research. It is being leveraged to construct surveys, synthesize data, conduct analysis, and write summaries of the results. While the promise is to create efficiencies and increase quality, the realit…

Can Agentic AI Match the Performance of Human Data Scientists?

2025-12-24 · An Luo, Jin Du, Fangqiao Tian, Xun Xian 외 arxiv

Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) have significantly automated data science workflows, b…