paper-with-me

Papers

ConDABench: Interactive Evaluation of Language Models for Data Analysis

2025-10-10 · Avik Dutta, Priyanshu Gupta, Hosein Hasanbeig, Rahul Pratap Singh, Harshit Nigam, Sumit Gulwani, Arjun Radhakrishna, Gustavo Soares, Ashish Tiwari arxiv

Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, essential to solving these complex tasks. Existing benchmarks for evaluating LLMs on data analysis tasks do not capture these complexities or provide first-class support for interactivity. We introduce ConDABench, a framework for generating conversational data analysis (ConDA) benchmarks and evaluating external tools on the generated benchmarks. \bench consists of (a) a multi-agent workflow for generating realistic benchmarks from articles describing insights gained from public datasets, (b) 1,420 ConDA problems generated using this workflow, and (c) an evaluation harness that, for the first time, makes it possible to systematically evaluate conversational data analysis tools on the generated ConDA problems. Evaluation of state-of-the-art LLMs on the benchmarks reveals that while the new generation of models are better at solving more instances, they are not necessarily better at solving tasks that require sustained, long-form engagement. ConDABench is an avenue for model builders to measure progress towards truly collaborative models that can complete complex interactive tasks.

📄 PDF Abstract BibTeX arXiv:2510.13835

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

2024-03-08 · Jinyang Li, Nan Huo, Yan Gao, Jiayi Shi 외

Interactive Data Analysis, the collaboration between humans and LLM agents, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic interactive logs for data anal…

BenchmarkingDecision MakingLanguage ModelingLanguage Modelling+1

IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents

2024-07-12 · Shrestha Mohanty, Negar Arabzadeh, Andrea Tupini, Yuxuan Sun 외

Seamless interaction between AI agents and humans using natural language remains a key goal in AI research. This paper addresses the challenges of developing interactive agents capable of understanding and executing grou…

Minecraft

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

2025-07-24 · Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim 외 arxiv

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not …

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

2025-05-21 · Xiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li 외

Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly focus on single-turn reasoning scenarios, leaving interactive tasks larg…

Attribute

PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis

2026-04-03 · Dexiang Li, Zhenning Che, Haijun Zhang, Dongliang Zhou 외 arxiv

Pavement condition assessment is essential for road safety and maintenance. Existing research has made significant progress. However, most studies focus on conventional computer vision tasks such as classification, detec…

Visual Question AnsweringSemantic SegmentationObject Detection