paper-with-me

홈 › Papers

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

2024-10-15 · Sijie Cheng, Kechen Fang, Yangyang Yu, Sicheng Zhou, Bohao Li, Ye Tian, Tingguang Li, Lei Han, Yang Liu

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric video understanding capabilities. To bridge the gap between MLLMs and low-level control in Embodied AI, we design four key interrelated tasks: video question-answering, hierarchy planning, visual grounding and reward modeling. To minimize manual annotation costs, we develop an automatic data generation pipeline based on the Ego4D dataset, leveraging the prior knowledge and multimodal capabilities of GPT-4o. Three human annotators then filter the generated data to ensure diversity and quality, resulting in the VidEgoThink benchmark. We conduct extensive experiments with three types of models: API-based MLLMs, open-source image-based MLLMs, and open-source video-based MLLMs. Experimental results indicate that all MLLMs, including GPT-4o, perform poorly across all tasks related to egocentric video understanding. These findings suggest that foundation models still require significant advancements to be effectively applied to first-person scenarios in Embodied AI. In conclusion, VidEgoThink reflects a research trend towards employing MLLMs for egocentric vision, akin to human capabilities, enabling active observation and interaction in the complex real-world environments.

📄 PDF Abstract BibTeX arXiv:2410.11623

Code (1)

adacheng/egothink pytorch

Tasks

Question AnsweringVideo Question AnsweringVideo UnderstandingVisual Grounding

Similar Papers 제목 키워드 기반

AMEGO: Active Memory from long EGOcentric videos

2024-09-17 · Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, Dima Damen

Egocentric videos provide a unique perspective into individuals' daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce AMEGO, a novel approach aimed at enhancing…

Video Understanding

X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding

2025-01-12 · Wenqi Zhou, Kai Cao, Hao Zheng, Xinyi Zheng 외

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activit…

Video Understanding

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

2025-03-12 · Haoyu Zhang, Qiaohui Chu, Meng Liu, Yunxiao Wang 외

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. Current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exoce…

Instruction FollowingVideo Understanding

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding

2024-06-19 · Alessandro Suglia, Claudio Greco, Katie Baker, Jose L. Part 외

AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person view videos, n…

Question AnsweringSpatial ReasoningVideo CaptioningVideo Question Answering+1

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

2025-06-11 · Benno Krojer, Mojtaba Komeili, Candace Ross, Quentin Garrido 외

Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visua…

Multiple-choice