paper-with-me

홈 › Papers

IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

2025-05-23 · Hanyu Li, Haoyu Liu, Tingyu Zhu, Tianyu Guo, Zeyu Zheng, Xiaotie Deng, Michael I. Jordan

Large Language Models (LLMs) show promise as data analysis agents, but existing benchmarks overlook the iterative nature of the field, where experts' decisions evolve with deeper insights of the dataset. To address this, we introduce IDA-Bench, a novel benchmark evaluating LLM agents in multi-round interactive scenarios. Derived from complex Kaggle notebooks, tasks are presented as sequential natural language instructions by an LLM-simulated user. Agent performance is judged by comparing its final numerical output to the human-derived baseline. Initial results show that even state-of-the-art coding agents (like Claude-3.7-thinking) succeed on < 50% of the tasks, highlighting limitations not evident in single-turn tests. This work underscores the need to improve LLMs' multi-round capabilities for building more reliable data analysis agents, highlighting the necessity of achieving a balance between instruction following and reasoning.

📄 PDF Abstract BibTeX arXiv:2505.18223

Code (1)

lhydave/ida-bench 공식 구현

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs

2025-08-13 · Kartikeya Badola, Jonathan Simon, Arian Hosseini, Sara Marie Mc Carthy 외 arxiv

Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights …

Instruction Following

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

2026-01-10 · Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang 외 arxiv

Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain largely autonomous, assuming fully specifi…

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

2025-09-20 · Junhao Chen, Jingbo Sun, Xiang Li, Haidong Xin 외 arxiv

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their i…

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

2026-05-26 · Jinzhao Li, Yinuo Chen, Wenxuan Song, Yijia Lei 외 arxiv

Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require proactive reasoning over continuous visual inputs. Existing benchmark…

Question Answering

CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling

2025-11-12 · Bichen Wang, Yixin Sun, Junzhe Wang, Hao Yang 외 arxiv

The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (LLMs) in this domain. Consequently, ther…