paper-with-me

Papers

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

2026-05-30 · Shengjun Zhang, Zhang Zhang, Simin Huang, Zhenyu Tang, Hanyang Wang, Chensheng Dai, Min Chen, Yifan Li, Yuxin Li, Yingjie Chen, Hao Liu, Chen Li, Jing Lyu, Yueqi Duan arxiv

Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fundamental gap persists between visually plausible video generation and the functional requirements of a world model, particularly in maintaining a stable and reasonable internal state over extended temporal horizons. While existing benchmarks primarily emphasize visual quality, motion coherence, and text-video alignment, they largely overlook memory, the core capability of a world model to preserve consistency across long-term horizons and complex interactions. To address this gap, we present \textbf{MBench}, a comprehensive benchmark dedicated to quantifying and evaluating the memory capability of video world models. We systematically decompose the memory capability of video world models into three hierarchical and complementary core dimensions: entity consistency, environment consistency, and causal consistency, which are further refined into 12 quantifiable sub-dimensions for comprehensive characterization of long-term memory. Our benchmark is built upon rigorously curated real-captured long videos, and evaluated by rule-based quantitative matrices and VLM to enable objective and comprehensive consistency assessment. Extensive evaluations of mainstream state-of-the-art video world models reveal critical systemic limitations of existing methods in long-term state retention, providing a standardized benchmark and clear research direction to advance the field.

📄 PDF Abstract BibTeX arXiv:2606.00793

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationVideo Alignment

Similar Papers 제목 키워드 기반

MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents

2025-06-20 · Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen 외

Recent works have highlighted the significance of memory mechanisms in LLM-based agents, which enable them to store observed information and adapt to dynamic environments. However, evaluating their memory capabilities st…

Diversity

MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

2024-06-20 · Xinyu Fang, Kangrui Mao, Haodong Duan, Xiangyu Zhao 외

The advent of large vision-language models (LVLMs) has spurred research into their applications in multi-modal contexts, particularly in video understanding. Traditional VideoQA benchmarks, despite providing quantitative…

FormVideo Understanding

FCMBench-Video: Benchmarking Document Video Intelligence

2026-04-28 · Runze Cui, Fangxin Shang, Yehui Yang, Qing Yang 외 arxiv

Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and evidence traceability matter. Compared with static document images, docume…

QMBench: A Research Level Benchmark for Quantum Materials Research

2025-12-19 · Yanzhen Wang, Yiyang Jiang, Diana Golovanova, Kamal Das 외 arxiv

We introduce QMBench, a comprehensive benchmark designed to evaluate the capability of large language model agents in quantum materials research. This specialized benchmark assesses the model's ability to apply condensed…

VMBench: A Benchmark for Perception-Aligned Video Motion Generation

2025-03-13 · Xinrang Ling, Chen Zhu, Meiqi Wu, Hangyu Li 외

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human…

Motion GenerationVideo Generation