paper-with-me

홈 › Papers

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

2026-05-22 · Darek Kleczek, Fuheng Zhao, Alexander W. Lee, Julien Tissier, Pawel Liskowski, Ugur Cetintemel, Anupam Datta arxiv

We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through \emph{latent world recovery}. AvalancheBench improves on existing benchmarks in three ways. First, it evaluates analytical understanding rather than pipeline completion: systems are scored on whether they recover the segments, drivers, temporal events, and relationships that explain the data, not merely on whether they execute a workflow or produce a plausible report. Second, it provides ground truth for goal-driven analytics by generating observations from a known latent world, enabling partial credit for incomplete but valid recoveries. Third, it exposes how early analytical mistakes propagate into later conclusions: missed segments, merged events, or wrong attributions can lead to systematically wrong recommendations. In this sense, AvalancheBench complements real-data benchmarks by providing a controlled setting for diagnosing whether agents recover the analytical structure behind enterprise data. On a first e-commerce use case, the strongest configuration of a leading coding agent recovers only 26\% of the rubric, with failures concentrated in generic customer segmentations and merged temporal events.

📄 PDF Abstract BibTeX arXiv:2605.24183

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DRBench: A Realistic Benchmark for Enterprise Deep Research

2025-09-30 · Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox 외 arxiv

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates …

Can LLMs Help You at Work? A Sandbox for Evaluating LLM Agents in Enterprise Environments

2025-10-31 · Harsh Vishwakarma, Ankush Agarwal, Ojas Patil, Chaitanya Devaguptapu 외 arxiv

Enterprise systems are crucial for enhancing productivity and decision-making among employees and customers. Integrating LLM based systems into enterprise systems enables intelligent automation, personalized experiences,…

Information Retrieval

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

2026-01-29 · Lakshya Gupta, Litao Li, Yizhe Liu, Sriram Ganapathi Subramanian 외 arxiv

Frontier large language models (LLMs) excel as autonomous agents in many domains, yet they remain untested in complex enterprise systems where hidden workflows create cascading effects across interconnected databases. Ex…

Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems

2025-11-18 · Sushant Mehta arxiv

Current agentic AI benchmarks predominantly evaluate task completion accuracy, while overlooking critical enterprise requirements such as cost-efficiency, reliability, and operational stability. Through systematic analys…

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

2026-05-09 · Tao Yu, Hao Wang, Changyu Li, Shenghua Chai 외 arxiv

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. How…