paper-with-me

Papers

EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents

2026-05-27 · Yunqi Liu, Tong Niu, Zitong Wang, Zhenlong Dai, Yuqi Qing, Weiqiang Wang, Jian Liu arxiv

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,045 egocentric-video-grounded tasks covering four daily scenarios, along with a user-agent-tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user within EgoBench to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight SOTA video-MLLM agents on EgoBench reveals a severe performance ceiling: the best model achieves only 30.62% accuracy in the best-performing scenario, averaging 19.43% across all four scenarios. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.

📄 PDF Abstract BibTeX arXiv:2605.27820

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

2025-03-12 · Haoyu Zhang, Qiaohui Chu, Meng Liu, Yunxiao Wang 외

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. Current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exoce…

Instruction FollowingVideo Understanding

LEGOBench: Scientific Leaderboard Generation Benchmark

2024-01-11 · Shruti Singh, Shoaib Alam, Husain Malwat, Mayank Singh

The ever-increasing volume of paper submissions makes it difficult to stay informed about the latest state-of-the-art research. To address this challenge, we introduce LEGOBench, a benchmark for evaluating systems that g…

DecoderLanguage ModelingLanguage Modelling

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

2026-05-17 · Zeyu Wang, Chang Liu, Eduardus Tjitrahardja, Yuntao Wang 외 arxiv

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remains largely overlooked. In this work, we i…

LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks

2026-02-28 · Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen 외 arxiv

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provid…

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

2026-04-13 · Weikun Peng, Denys Iliash, Manolis Savva arxiv

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling…