paper-with-me

Papers

Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents

2025-06-10 · Irene Testini, José Hernández-Orallo, Lorenzo Pacchiardi

Data science aims to extract insights from data to support decision-making processes. Recently, Large Language Models (LLMs) are increasingly used as assistants for data science, by suggesting ideas, techniques and small code snippets, or for the interpretation of results and reporting. Proper automation of some data-science activities is now promised by the rise of LLM agents, i.e., AI systems powered by an LLM equipped with additional affordances--such as code execution and knowledge bases--that can perform self-directed actions and interact with digital environments. In this paper, we survey the evaluation of LLM assistants and agents for data science. We find (1) a dominant focus on a small subset of goal-oriented activities, largely ignoring data management and exploratory activities; (2) a concentration on pure assistance or fully autonomous agents, without considering intermediate levels of human-AI collaboration; and (3) an emphasis on human substitution, therefore neglecting the possibility of higher levels of automation thanks to task transformation.

📄 PDF Abstract BibTeX arXiv:2506.08800

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

A Survey of Open Source Automation Tools for Data Science Predictions

2022-08-24 · Nicholas Hoell

We present an expository overview of technical and cultural challenges to the development and adoption of automation at various stages in the data science prediction lifecycle, restricting focus to supervised learning wi…

Converging Measures and an Emergent Model: A Meta-Analysis of Human-Automation Trust Questionnaires

2023-03-24 · Yosef S. Razin, Karen M. Feigh

A significant challenge to measuring human-automation trust is the amount of construct proliferation, models, and questionnaires with highly variable validation. However, all agree that trust is a crucial element of tech…

Common Sense ReasoningSurvey

Measuring Readability of Polish Texts: Baseline Experiments

2014-05-01 · LREC 2014 5 · Bartosz Broda, Bart{\l}omiej Nito{\'n}, W{\l}odzimierz Gruszczy{\'n}ski, Maciej Ogrodniczuk

Measuring readability of a text is the first sensible step to its simplification. In this paper we present an overview of the most common approaches to automatic measuring of readability. Of the described ones, we implem…

Language Modelling

Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions

2025-03-12 · Mourad Gridach, Jay Nanavati, Khaldoun Zine El Abidine, Lenon Mendes 외

The integration of Agentic AI into scientific discovery marks a new frontier in research automation. These AI systems, capable of reasoning, planning, and autonomous decision-making, are transforming how scientists perfo…

Decision Makingscientific discovery

From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery

2025-05-19 · Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang 외

Large Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and h…

Navigatescientific discoverySurvey