paper-with-me

Papers

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

2025-01-09 · CVPR 2025 1 · Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, Lidong Bing

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code are available at https://github.com/Rh-Dang/ECBench.

📄 PDF Abstract BibTeX arXiv:2501.05031

Code (1)

rh-dang/ecbench 공식 구현

Tasks

FairnessHallucinationQuestion AnsweringVideo Question Answering

Similar Papers 제목 키워드 기반

MM-Ego: Towards Building Egocentric Multimodal LLMs

2024-10-09 · Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen 외

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric …

Video Understanding

EgoM2P: Egocentric Multimodal Multitask Pretraining

2025-06-09 · Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao 외

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction. These capabilities en…

Depth EstimationGaze PredictionMonocular Depth Estimation

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

2026-06-18 · Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le 외 arxiv

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …

Representation LearningAction SegmentationAction RecognitionVideo Retrieval

Aria-NeRF: Multimodal Egocentric View Synthesis

2023-11-11 · Jiankai Sun, Jianing Qiu, Chuanyang Zheng, John Tucker 외

We seek to accelerate research in developing rich, multimodal scene models trained from egocentric data, based on differentiable volumetric ray-tracing inspired by Neural Radiance Fields (NeRFs). The construction of a Ne…

NeRF

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

2026-02-15 · Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about…

Causal Inference