paper-with-me

홈 › Papers

MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models

2026-05-28 · Tianzhuo Yang, Zihan Shen, Zirui Mi, Zhaoyi Zhang, Jiayi Zhou, Jiaming Ji, Juntao Dai, Jiawei Chen, Boyuan Chen, Yaodong Yang arxiv

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks largely emphasize visual fidelity, leaving unclear whether predicted futures are physically plausible, faithful to commanded actions, and calibrated to failure when actions should not succeed. We introduce \textsc{MiraBench}, a hierarchical benchmark that defines \emph{action-conditioned reliability} as a core evaluation target for robotic world models. MiraBench decomposes this target into three progressively demanding levels: \emph{Physics Adherence}, which evaluates reference-free physical consistency; \emph{Action-Following Fidelity}, which measures whether predictions respect task-relevant action inputs; and \emph{Optimism Bias Detection}, which probes the tendency to predict successful outcomes under failure-inducing actions. To support this evaluation, we curate a human-annotated corpus with over 16,000 judgments across tasks, failure categories, and leading world models. We evaluate 12 representative model configurations spanning vector-conditioned robotic world models, text-conditioned generative world models, open-weight systems, closed-source systems, and multiple model scales. Across this broad model landscape, MiraBench reveals three central findings: visual fidelity is a poor proxy for action fidelity; increasing model scale does not reliably improve action following; and optimism bias is pervasive across current systems. By shifting evaluation from appearance to action-conditioned reliability, MiraBench provides a diagnostic foundation for assessing and improving robotic world models as faithful simulators.

📄 PDF Abstract BibTeX arXiv:2605.29360

Code (0)

등록된 구현이 없습니다.

Tasks

Bias Detection

Similar Papers 제목 키워드 기반

MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

2024-07-08 · Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan 외

Sora's high-motion intensity and long consistent videos have significantly impacted the field of video generation, attracting unprecedented attention. However, existing publicly available datasets are inadequate for gene…

Video AlignmentVideo Generation

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

2025-10-19 · Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri 외 arxiv

Cardiac catheterization remains a cornerstone of minimally invasive interventions, yet it continues to rely heavily on manual operation. Despite advances in robotic platforms, existing systems are predominantly follow-le…

How Should a Robot Configure Its Laser Scanner for Inspection?

2026-06-19 · Zhiling Chen, David Gorsich, Matthew P. Castanier, Yang Zhang 외 arxiv

Robotic inspection relies on accurate sensing to acquire high-fidelity geometric measurements for defect detection and metrology. While prior work has focused on robot motion and viewpoint planning, how to configure sens…

World Models for Robotic Manipulation: A Survey

2026-05-27 · Fangyuan Wang, Ziyuan Wang, Guorui Pei, Mengshi Zhang 외 arxiv

Robotic manipulation depends on the ability to anticipate how actions reshape objects, contacts, and scene geometry before execution. Learned world models provide this capability by predicting task-relevant future evolut…

Scalable Policy Evaluation with Video World Models

2025-11-14 · Wei-Cheng Tseng, Jinwei Gu, Qinsheng Zhang, Hanzi Mao 외 arxiv

Training generalist policies for robotic manipulation has shown great promise, as they enable language-conditioned, multi-task behaviors across diverse scenarios. However, evaluating these policies remains difficult beca…

Video Generation