paper-with-me

홈 › Papers

OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios

2026-07-16 · Chengyu Shen, Yujie Fu, Gangtao Xin, Yanheng Hou, Wenlong Fei, Guojie Zhu, Jiawei Li, Hongcheng Gao, Runming He, Zhen Hao Wong, Meiyi Qiang, Hao Liang, Zhao Cao, Hao Jiang, Chong Chen, Wentao Zhang arxiv

Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction. However, existing agent benchmarks often focus on limited scenarios, tool ecosystems, or interaction formats, making it difficult to systematically characterize model capabilities across heterogeneous application settings. We introduce OmniaBench, a benchmark for evaluating general agents across diverse scenarios with explicit state spaces. We derive application-oriented scenario knowledge from app stores, product documents, industry resources, Web retrieval, and human refinement, forming a hierarchical taxonomy that spans ToC, ToB and ToE with 90 level-1 and 354 level-2 domains. Based on this taxonomy, we construct executable environments and synthesize single-turn and multi-turn tasks through four complementary routes: DAG, DAG-S, Solver, and Program. OmniaBench further introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis. The resulting dataset contains 1,431 tasks, together with a challenging subset of 644 tasks designed to reduce evaluation cost and mitigate potential contamination of the full set after public release. The bench presents substantial challenges to current frontier models, with even Claude-Sonnet-5 and GPT-5.6-Sol achieving Overall Pass@1 scores of only 58.54 and 57.14, respectively. Further analyses reveal clear differences across domains and capabilities, as well as persistent limitations in planning, constraint maintenance, and adaptive correction. OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents.

📄 PDF Abstract BibTeX arXiv:2607.14989

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

2026-07-25 · Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu 외 arxiv

Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings. We frame this as full-scenario agentic scaling and present Agent…

Reinforcement Learning

Benchmarking Mobile Device Control Agents across Diverse Configurations

2024-04-25 · Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm 외

Mobile device control agents can largely enhance user interactions and productivity by automating daily tasks. However, despite growing interest in developing practical agents, the absence of a commonly adopted benchmark…

BenchmarkingImitation Learning

Towards Evaluating Generalist Agents: An Automated Benchmark in Open World

2023-10-12 · Xinyue Zheng, Haowei Lin, Kaichen He, ZiHao Wang 외

Evaluating generalist agents presents significant challenges due to their wide-ranging abilities and the limitations of current benchmarks in assessing true generalization. We introduce the Minecraft Universe (MCU), a fu…

BenchmarkingDiversityLanguage ModelingLanguage Modelling+2

Benchmarking the Generality of Vision-Language-Action Models

2025-12-12 · Pranav Guruprasad, Sudipta Chowdhury, Harsh Sikka, Mridul Sharma 외 arxiv

Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices remain fragmented across isolated benchma…

Spatial ReasoningVisual Grounding

CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions

2025-05-24 · Kung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal 외

While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platforms. Existing benchmarks often lack fideli…

Benchmarking