paper-with-me

Papers

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

2025-06-10 · Wendong Bu, Yang Wu, Qifan Yu, Minghe Gao, Bingchen Miao, Zhenkui Zhang, Kaihang Pan, Yunfei Li, Mengze Li, Wei Ji, Juncheng Li, Siliang Tang, Yueting Zhuang

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity, extensive manual annotation with limited scenarios, and a lack of multidimensional evaluation. In response to these challenges, we introduce OmniBench, a self-generating, cross-platform, graph-based benchmark with an automated pipeline for synthesizing tasks of controllable complexity through subtask composition. To evaluate the diverse capabilities of virtual agents on the graph, we further present OmniEval, a multidimensional evaluation framework that includes subtask-level evaluation, graph-based metrics, and comprehensive tests across 10 capabilities. Our synthesized dataset contains 36k graph-structured tasks across 20 scenarios, achieving a 91\% human acceptance rate. Training on our graph-structured data shows that it can more efficiently guide agents compared to manually annotated data. We conduct multidimensional evaluations for various open-source and closed-source models, revealing their performance across various capabilities and paving the way for future advancements. Our project is available at https://omni-bench.github.io/.

📄 PDF Abstract BibTeX arXiv:2506.08933

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Omnibenchmark (alpha) for continuous and open benchmarking in bioinformatics

2024-09-25 · Izaskun Mallona, Almut Luetge, Ben Carrillo, Daniel Incicau 외

Benchmarking in bioinformatics is a process of designing, running and disseminating rigorous performance evaluations of methods (software). Benchmarking systems facilitate the benchmarking process by providing an entrypo…

Benchmarking

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

2026-03-19 · Keda Tao, Yuhua Zheng, Jia Xu, Wenjie Du 외 arxiv

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips rangi…

OmniBench: Towards The Future of Universal Omni-Language Models

2024-09-23 · Yizhi Li, Ge Zhang, Yinghao Ma, Ruibin Yuan 외

Recent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We in…

Instruction Following

Virtual Mouse And Assistant: A Technological Revolution Of Artificial Intelligence

2023-03-11 · Jagbeer Singh, Yash Goel, Shubhi Jain, Shiva Yadav

The purpose of this paper is to enhance the performance of the virtual assistant. So, what exactly is a virtual assistant. Application software, often called virtual assistants, also known as AI assistants or digital ass…

Scheduling

AI in a vat: Fundamental limits of efficient world modelling for agent sandboxing and interpretability

2025-04-06 · Fernando Rosas, Alexander Boyd, Manuel Baltieri

Recent work proposes using world models to generate controlled virtual environments in which AI agents can be tested before deployment to ensure their reliability and safety. However, accurate world models often have hig…

AI Agent