paper-with-me

Papers

RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

2025-05-04 · Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Yibo Yan, Hanqian Li, Linghao Zhang, Shikang Wang, Yixin Liu, Hanbo Zhang, Ying Ma, Xuming Hu

Multimodal Large Language Models (MLLMs) increasingly excel at perception, understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RTV-Bench, a fine-grained benchmark for MLLM real-time video analysis. RTV-Bench uses three key principles: (1) Multi-Timestamp Question Answering (MTQA), where answers evolve with scene changes; (2) Hierarchical Question Structure, combining basic and advanced queries; and (3) Multi-dimensional Evaluation, assessing the ability of continuous perception, understanding, and reasoning. RTV-Bench contains 552 diverse videos (167.2 hours) and 4,631 high-quality QA pairs. We evaluated leading MLLMs, including proprietary (GPT-4o, Gemini 2.0), open-source offline (Qwen2.5-VL, VideoLLaMA3), and open-source real-time (VITA-1.5, InternLM-XComposer2.5-OmniLive) models. Experiment results show open-source real-time models largely outperform offline ones but still trail top proprietary models. Our analysis also reveals that larger model size or higher frame sampling rates do not significantly boost RTV-Bench performance, sometimes causing slight decreases. This underscores the need for better model architectures optimized for video stream processing and long sequences to advance real-time video analysis with MLLMs. Our benchmark toolkit is available at: https://github.com/LJungang/RTV-Bench.

📄 PDF Abstract BibTeX arXiv:2505.02064

Code (1)

ljungang/rtv-bench 공식 구현 pytorch

Tasks

BenchmarkingQuestion Answering

Similar Papers 제목 키워드 기반

Benchmarking Visual State Tracking in Multimodal Video Understanding

2026-06-02 · Sihyun Yu, Nanye Ma, Pinzhi Huang, Hyunseok Lee 외 arxiv

Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacity for visual state tracking is fundamental to video understanding, y…

FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs

2025-03-27 · CVPR 2025 1 · Xiaoqin Wang, Xusen Ma, Xianxu Hou, Meidan Ding 외

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in various tasks. However, effectively evaluating these MLLMs on face perception remains largely unexplored. To address this gap, we intr…

AttributeBenchmarkingQuestion AnsweringVisual Question Answering+1

Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models

2025-08-06 · Zhaochen Liu, Kaiwen Gao, Shuyi Liang, Bin Xiao 외 arxiv

Occlusion perception, a critical foundation for human-level spatial understanding, embodies the challenge of integrating visual recognition and reasoning. Though multimodal large language models (MLLMs) have demonstrated…

Visual Question Answering

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

2025-08-18 · Ashish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to…

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

2026-02-15 · Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about…

Causal Inference