paper-with-me

홈 › Papers

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

2026-05-28 · Kewei Xu, Xiaoben Lu, Shuofei Qiao, Zihan Ding, Haoming Xu, Lei Liang, Ningyu Zhang arxiv

Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long horizons untested. We introduce LongDS, a benchmark for long-horizon, multi-turn data analysis where agents must maintain, update, restore, and compose evolving analytical states. LongDS comprises 68 tasks constructed from real-world Kaggle notebooks, spanning 2,225 turns across six domains including Geoscience, Business, and Education. Tasks are designed around state-evolution patterns (e.g., counterfactual perturbation, rollback, multi-state composition), with an average dependency span of 11.3 turns. Evaluating five state-of-the-art models, we find that the best model reaches only 48.45% average accuracy, performance drops nearly 47 points from early to late turns, and long-horizon errors account for 52%--69% of failures. Further analysis shows that additional agent steps do not necessarily improve performance, suggesting that the key bottleneck is maintaining a correct analytical state rather than increasing interaction budget. We release LongDS to support research on reliable long-horizon agentic data analysis. Code and data are released at https://github.com/zjunlp/DataMind.

📄 PDF Abstract BibTeX arXiv:2605.30434

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

2026-04-13 · Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun, Haorui Wang 외 arxiv

Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, interdependent action sequences. Despite rapid progress in agentic sys…

Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework

2026-05-02 · Mukund Pandey arxiv

Existing evaluation frameworks for large language models -- including HELM, MT-Bench, AgentBench, and BIG-bench -- are designed for controlled, single-session, lab-scale settings. They do not address the evaluation chall…

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

2026-04-02 · Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu 외 arxiv

Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses. Existing trajectory-le…

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

2026-08-28 · Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao 외 arxiv

Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-…

PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems

2026-06-21 · Jiayu Liu, Qihan Lin, Cheng Qian, Rui Wang 외 arxiv

LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existin…