paper-with-me

Papers

WorldBench: Benchmarking Physical Understanding of World Models by Isolating Physics Concepts

2026-01-29 · Rishi Upadhyay, Howard Zhang, Jim Solomon, Ayush Agrawal, Pranay Boreddy, Shruti Satya Narayana, Yunhao Ba, Alex Wong, Celso M de Melo, Achuta Kadambi arxiv

Recent advances in generative foundational models, often termed "world models," have propelled interest in applying them to critical tasks like robotic planning and autonomous system training. For reliable deployment, these models must exhibit high physical fidelity, accurately simulating real-world dynamics. Existing physics-based video benchmarks, however, suffer from entanglement, where a single test simultaneously evaluates multiple physical laws and concepts, fundamentally limiting their diagnostic capability. We introduce WorldBench, a novel video-based benchmark specifically designed for concept-specific, disentangled evaluation, allowing us to rigorously isolate and assess understanding of a single physical concept or law at a time. To make WorldBench comprehensive, we design benchmarks at two different levels: 1) an evaluation of intuitive physical understanding with higher level concepts such as object permanence or scale/perspective, and 2) an evaluation of low-level physical constants and material properties such as friction coefficients or fluid viscosity, allowing to measure excatly how far from reality generated videos are. When SOTA video-based world models are evaluated on WorldBench, we find specific patterns of failure in particular physics concepts, with all tested models lacking the physical consistency required to generate reliable real-world interactions. Through its concept-specific evaluation, WorldBench offers a more nuanced and scalable framework for rigorously evaluating the physical reasoning capabilities of video generation and world models, paving the way for more robust and generalizable world-model-driven learning.

📄 PDF Abstract BibTeX arXiv:2601.21282

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

2026-06-04 · Yida Yin, Harish Krishnakumar, Chung Peng Lee, Boya Zeng 외 arxiv

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual diversity needed to handle open-ended v…

Multimodal Reasoning

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

2025-11-25 · Yiting Lu, Wei Luo, Peiyan Tu, Haoran Li 외 arxiv

World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct realistic, dynamic, and physically consiste…

Autonomous Driving

"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

2025-07-17 · Jing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan 외 arxiv

Video generation models have achieved remarkable progress in creating high-quality, photorealistic content. However, their ability to accurately simulate physical phenomena remains a critical and unresolved challenge. Th…

Text-to-Video Generation

MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation

2026-02-28 · Rongsheng Wang, Minghao Wu, Hongru Zhou, Zhihan Yu 외 arxiv

Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds gr…

Instruction FollowingVideo GenerationDrug Discovery

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

2026-03-23 · Meiqi Wu, Zhixin Cai, Fufangchen Zhao, Xiaokun Feng 외 arxiv

Video--based world models have emerged along two dominant paradigms: video generation and 3D reconstruction. However, existing evaluation benchmarks either focus narrowly on visual fidelity and text--video alignment for …

3D ReconstructionVideo GenerationVideo Alignment