paper-with-me

홈 › Papers

Benchmarking Data Science Agents

2024-02-27 · Yuge Zhang, Qiyang Jiang, Xingyu Han, Nan Chen, Yuqing Yang, Kan Ren

In the era of data-driven decision-making, the complexity of data analysis necessitates advanced expertise and tools of data science, presenting significant challenges even for specialists. Large Language Models (LLMs) have emerged as promising aids as data science agents, assisting humans in data analysis and processing. Yet their practical efficacy remains constrained by the varied demands of real-world applications and complicated analytical process. In this paper, we introduce DSEval -- a novel evaluation paradigm, as well as a series of innovative benchmarks tailored for assessing the performance of these agents throughout the entire data science lifecycle. Incorporating a novel bootstrapped annotation method, we streamline dataset preparation, improve the evaluation coverage, and expand benchmarking comprehensiveness. Our findings uncover prevalent obstacles and provide critical insights to inform future advancements in the field.

📄 PDF Abstract BibTeX arXiv:2402.17168

Code (1)

metacopilot/dseval 공식 구현

Tasks

BenchmarkingCode GenerationDecision Making

Similar Papers 제목 키워드 기반

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

2025-10-24 · Jonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket 외 arxiv

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many su…

DSBC : Data Science task Benchmarking with Context engineering

2025-07-31 · Ram Mohan Rao Kadiyala, Siddhant Gupta, Jebish Purbey, Giulio Martini 외 arxiv

Recent advances in large language models (LLMs) have significantly impacted data science workflows, giving rise to specialized data science agents designed to automate analytical tasks. Despite rapid adoption, systematic…

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

2026-02-11 · Bang Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage 외 arxiv

The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to repro…

CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories

2025-02-10 · Yijia Xiao, Runhui Wang, Luyang Kong, Davor Golac 외

The increasing complexity of computer science research projects demands more effective tools for deploying code repositories. Large Language Models (LLMs), such as Anthropic Claude and Meta Llama, have demonstrated signi…

Benchmarking

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

2026-03-19 · An Luo, Jin Du, Xun Xian, Robert Specht 외 arxiv

Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large language models (LLMs) and artificial intelligence (AI) agents have significa…