paper-with-me

Papers

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

2026-06-25 · Haoyu Chen, Kaichen Zhou, Hang Hua, Kaile Zhang, Jingwen Qian, Wufei Ma, Haonan Chen, Chunjiang Liu, Yizhou Zhao, Xiaoyuan Wang, Weiyue Li, Alan Yuille, Paul Pu Liang, Yilun Du arxiv

Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few that force objects out of view evaluate static scenes where nothing changes during occlusion. To bridge this gap, we introduce MemoBench, a diagnostic benchmark built around the disappear-and-reappear paradigm in dynamically changing environments: a target object undergoes a physical process, disappears from view, and must be correctly recovered in its updated state upon reappearance. We curate 360 ground-truth clips spanning synthetic and real-world scenes, and design an evaluation suite combining automated metrics with VQA-based assessment across four diagnostic pillars. Evaluation of eight state-of-the-art models reveals key insights and open challenges regarding memory consistency under the disappear-and-reappear paradigm.

📄 PDF Abstract BibTeX arXiv:2606.27537

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

AllSim: Simulating and Benchmarking Resource Allocation Policies in Multi-User Systems

2023-09-26 · NeurIPS 2023 11

Numerous real-world systems, ranging from healthcare to energy grids, involve users competing for finite and potentially scarce resources. Designing policies for resource allocation in such real-world systems is challeng…

Discrete models of continuous behavior of collective adaptive systems

2022-04-26 · Peter Fettke, Wolfgang Reisig

Artificial ants are "small" units, moving autonomously on a shared, dynamically changing "space", directly or indirectly exchanging some kind of information. Artificial ants are frequently conceived as a paradigm for col…

HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments

2024-01-23 · Qinhong Zhou, Sunli Chen, Yisong Wang, Haozhe Xu 외

Recent advances in high-fidelity virtual environments serve as one of the major driving forces for building intelligent embodied agents to perceive, reason and interact with the physical world. Typically, these environme…

Common Sense ReasoningDecision MakingReinforcement Learning (RL)

Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views

2023-02-23 · Katerina Margatina, Shuai Wang, Yogarshi Vyas, Neha Anna John 외

Temporal concept drift refers to the problem of data changing over time. In NLP, that would entail that language (e.g. new expressions, meaning shifts) and factual knowledge (e.g. new concepts, updated facts) evolve over…

Benchmarking

NEST: A Neuromodulated Small-world Hypergraph Trajectory Prediction Model for Autonomous Driving

2024-12-16 · Chengyue Wang, Haicheng Liao, Bonan Wang, Yanchen Guan 외

Accurate trajectory prediction is essential for the safety and efficiency of autonomous driving. Traditional models often struggle with real-time processing, capturing non-linearity and uncertainty in traffic environment…

Autonomous DrivingPredictionTrajectory Prediction