paper-with-me

Papers

EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports

2026-04-14 · Jianzhe Ma, Zhonghao Cao, Shangkui Chen, Yichen Xu, Wenxuan Wang, Qin Jin arxiv

While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on daily activities, yet lack a rigorous testbed for evaluating fast, rule-bound reasoning in virtual scenarios. To fill this gap, we introduce EgoEsportsQA, a pioneering video question-answering (QA) benchmark for grounding perception and reasoning in expert esports knowledge. We curate 1,745 high-quality QA pairs from professional matches across 3 first-person shooter games via a scalable six-stage pipeline. These questions are structured into a two-dimensional decoupled taxonomy: 11 sub-tasks in the cognitive capability dimension (covering perception and reasoning levels) and 6 sub-tasks in the esports knowledge dimension. Comprehensive evaluations of state-of-the-art Video-LLMs reveal that current models still fail to achieve satisfactory performance, with the best model only 71.58%. The results expose notable gaps across both axes: models exhibit stronger capabilities in basic visual perception than in deep tactical reasoning, and they grasp overall macro-progression better than fine-grained micro-operations. Extensive ablation experiments demonstrate the intrinsic weaknesses of current Video-LLM architectures. Further analysis suggests that our dataset not only reveals the connections between real-world and virtual egocentric domains, but also offers guidance for optimizing downstream esports applications, thereby fostering the future advancement of Video-LLMs in various egocentric environments.

📄 PDF Abstract BibTeX arXiv:2604.12320

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

2026-05-19 · Yang Dai, Dian Jiao, Tianwei Lin, Wenqiao Zhang arxiv

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, trac…

Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

2026-06-13 · Yun Wang, Junbin Xiao, Han Lyu, Yifan Wang 외 arxiv

We introduce UCS-Bench, a dataset spanning 170+ hours of egocentric visual observations with 8.1K+ timestamped questions for diagnosing User-Centric Continual Spatial intelligence in egocentric video streams. UCS-Bench t…

Spatial Reasoning

AMEGO: Active Memory from long EGOcentric videos

2024-09-17 · Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, Dima Damen

Egocentric videos provide a unique perspective into individuals' daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce AMEGO, a novel approach aimed at enhancing…

Video Understanding

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

2026-02-15 · Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about…

Causal Inference