paper-with-me

홈 › Papers

EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs

2025-11-23 · Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Xiangyang Ji arxiv

Multimodal large language models (MLLMs) have made significant advancements in event-based vision, yet the comprehensive evaluation of their capabilities within a unified benchmark remains largely unexplored. In this work, we introduce EventBench, a benchmark that offers eight diverse task metrics together with a large-scale event stream dataset. EventBench differs from existing event-based benchmarks in four key aspects: (1) openness in accessibility, releasing all raw event streams and task instructions across eight evaluation metrics; (2) diversity in task coverage, spanning understanding, recognition, and spatial reasoning tasks for comprehensive capability assessment; (3) integration in spatial dimensions, pioneering the design of 3D spatial reasoning tasks for event-based MLLMs; and (4) scale in data volume, with an accompanying training set of over one million event-text pairs supporting large-scale training and evaluation. Using EventBench, we evaluate state-of-the-art closed-source models such as GPT-5 and Gemini-2.5 Pro, leading open-source models including Qwen2.5-VL and InternVL3, and event-based MLLMs such as EventGPT that directly process raw event streams. Extensive evaluation reveals that while current event-based MLLMs demonstrate strong performance in event stream understanding, they continue to struggle with fine-grained recognition and spatial reasoning.

📄 PDF Abstract BibTeX arXiv:2511.18448

Code (0)

등록된 구현이 없습니다.

Tasks

Event-based visionSpatial Reasoning

Similar Papers 제목 키워드 기반

EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models

2025-11-24 · Wenhao Xu, Xin Dong, Yue Li, Haoyuan Shi 외 arxiv

Video large language models have demonstrated strong video understanding capabilities but suffer from high inference costs due to the massive number of tokens in long videos. Inspired by event-based vision, we propose an…

Event-based vision

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

2026-07-17 · Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou 외 arxiv

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets a…

Reinforcement Learning

Question-Answering Dense Video Events

2024-09-06 · Hangyu Qin, Junbin Xiao, Angela Yao

This paper presents question-answering on dense video events, a novel task that answers and grounds dense-event questions in long videos, thus challenging MLLMs to faithfully comprehend and reason about multiple events o…

BenchmarkingQuestion AnsweringZero-Shot Video Question Answer

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

2026-07-16 · Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang 외 arxiv

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general visio…

Visual Question Answering

Benchmarking Visual State Tracking in Multimodal Video Understanding

2026-06-02 · Sihyun Yu, Nanye Ma, Pinzhi Huang, Hyunseok Lee 외 arxiv

Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, y…