paper-with-me

홈 › Papers

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

2026-07-28 · Hao Liang, Meiyi Qiang, Sizhe Qiu, Linzhuang Sun, Wentao Zhang arxiv

Enterprise agents often need to integrate heterogeneous knowledge sources: documents for narrative facts, tables for computation, and dependency graphs for file relationships. Existing benchmarks typically evaluate retrieval or tool use without distinguishing whether an agent first selects the appropriate knowledge sources. We introduce WorkSurface-Bench, a benchmark for evaluating this capability as surface routing. It contains 1,151 atomic tasks derived from persona-scoped Workspace-Bench-Lite workspaces, spanning document, table, graph, and cross-surface questions. Its reference answers are auditable: table answers are reproduced through executed DuckDB queries, document answers are grounded in verified text spans, and graph answers are traced to source dependency annotations. We evaluate four model backbones across six controlled agent settings, yielding 27,624 protocol-error-free trajectories. Under gold-constrained tool access, agents achieve 98.7-99.8 Route F1, while Answer remains only 56.1-75.3 percent, showing that correct surface selection is necessary but insufficient for task completion. Matched interventions further show that surface hints improve Answer for three of four models, whereas removing irrelevant tools primarily improves routing and efficiency. In an independent three-annotator audit, all 200 sampled tasks pass all six quality criteria by majority vote, with 192 receiving unanimous judgments on every criterion. We release the dataset, construction pipeline, scoring code, and agent harness at https://github.com/haolpku/WorkSurface-Bench.

📄 PDF Abstract BibTeX arXiv:2607.25765

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows

2026-05-09 · Tao Yu, Hao Wang, Changyu Li, Shenghua Chai 외 arxiv

Large language model (LLM) agents are increasingly expected to operate in enterprise environments, where work is distributed across specialized roles, permission-controlled systems, and cross-departmental procedures. How…

EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions

2026-06-22 · Jincheng Zhong, Weizhi Wang, Che Jiang, Kai Tian 외 arxiv

Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark constructed from prop…

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

2026-04-23 · Wenjie Fu, Xiaoting Qin, Jue Zhang, Qingwei Lin 외 arxiv

Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user's behalf, also creates new risks for sensitive information leakage.…

Benchmarking Agents in Insurance Underwriting Environments

2026-01-31 · Amanda Dsouza, Ramya Ramakrishnan, Charles Dickens, Bhavishya Pohani 외 arxiv

As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemphasize open-domains such as code, use nar…

The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents

2026-09-09 · Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida 외 arxiv

LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no existing substitute provides ground truth. We present the Era by Eon Benchmark for evaluating LLM agents that use enterp…