paper-with-me

홈 › Papers

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

2025-02-03 · Harshith Padigela, Chintan Shah, Dinkar Juyal

In this report, we present ML-Dev-Bench, a benchmark aimed at testing agentic capabilities on applied Machine Learning development tasks. While existing benchmarks focus on isolated coding tasks or Kaggle-style competitions, ML-Dev-Bench tests agents' ability to handle the full complexity of ML development workflows. The benchmark assesses performance across critical aspects including dataset handling, model training, improving existing models, debugging, and API integration with popular ML tools. We evaluate three agents - ReAct, Openhands, and AIDE - on a diverse set of 30 tasks, providing insights into their strengths and limitations in handling practical ML development challenges. We open source the benchmark for the benefit of the community at \href{https://github.com/ml-dev-bench/ml-dev-bench}{https://github.com/ml-dev-bench/ml-dev-bench}.

📄 PDF Abstract BibTeX arXiv:2502.00964

Code (1)

ml-dev-bench/ml-dev-bench 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Accountable Agents in Software Engineering: An Analysis of Terms of Service and a Research Roadmap

2026-05-06 · Christoph Treude arxiv

AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities a…

ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

2025-05-26 · Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding 외

Large Language Models (LLMs) have extended their impact beyond Natural Language Processing, substantially fostering the development of interdisciplinary research. Recently, various LLM-based agents have been developed to…

Astronomyscientific discovery

HeurekaBench: A Benchmarking Framework for AI Co-scientist

2026-01-04 · Siba Smarak Panigrahi, Jovana Videnović, Maria Brbić arxiv

LLM-based reasoning models have enabled the development of agentic systems that act as co-scientists, assisting in multi-step scientific analysis. However, evaluating these systems is challenging, as it requires realisti…

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows

2026-04-17 · Jize Wang, Xuanxuan Liu, Yining Li, Songyang Zhang 외 arxiv

The development of general-purpose agents requires a shift from executing simple instructions to completing complex, real-world productivity workflows. However, current tool-use benchmarks remain misaligned with real-wor…

SastBench: A Benchmark for Testing Agentic SAST Triage

2026-01-06 · Jake Feiglin, Guy Dar arxiv

SAST (Static Application Security Testing) tools are among the most widely used techniques in defensive cybersecurity, employed by commercial and non-commercial organizations to identify potential vulnerabilities in soft…