paper-with-me

Papers

ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models

2026-05-09 · Haotian Xue, Yipu Chen, Liqian Ma, Zelin Zhao, Lama Moukheiber, Yuchen Zhu, Yongxin Chen arxiv

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering only limited coverage of the rich physical interactions required for generalized world understanding. We introduce ACWM-Phys, a new benchmark for evaluating action-conditioned prediction under diverse physical dynamics in a clean, controllable simulation environment with a carefully designed action space. ACWM-Phys contains training and evaluation data spanning rigid-body dynamics, kinematics, deformable-object interactions, and particle dynamics. To evaluate both interpolation and generalization, we design in-distribution and out-of-distribution protocols with controlled shifts in interaction patterns or scene configurations. By building the benchmark in a fully controllable simulator, ACWM-Phys enables precise data collection, reproducible evaluation, and systematic analysis of model capabilities for physically grounded world modeling. Through systematic experiments on ACWM-DiT, we find that OoD generalization depends not only on the physical regime but also on effective task complexity: models generalize well on visually simple, low-dimensional interactions with clear geometric structure, but suffer larger drops on deformable contacts, high-dimensional control, and complex articulated motion. This suggests that the model still relies heavily on visual appearance patterns instead of fully learning the underlying physics. Ablations show that cross-attention improves high-dimensional action conditioning, causal VAEs outperform frame-wise encoders, and larger action spaces are harder to model but can improve generalization by providing richer control signals. These findings guide the design of physically grounded world models.

📄 PDF Abstract BibTeX arXiv:2605.08567

Code (0)

등록된 구현이 없습니다.

Tasks

Video Prediction

Similar Papers 제목 키워드 기반

Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

2026-07-10 · Yufan Wei, Kun Zhou, Lingjun Mao, Zijun Zhang 외 arxiv

Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning…

Contrastive LearningData Augmentation

Occlusion-Robust Multi-Sensory Posture Estimation in Physical Human-Robot Interaction

2022-08-12 · Amir Yazdani, Roya Sabbagh Novin, Andrew Merryweather, Tucker Hermans

3D posture estimation is important in analyzing and improving ergonomics in physical human-robot interaction and reducing the risk of musculoskeletal disorders. Vision-based posture estimation approaches are prone to sen…

Investigating the impact of free energy based behavior on human in human-agent interaction

2022-01-25 · Kazuya Horibe, Yuanxiang Fan, Yutaka Nakamura, Hiroshi Ishiguro

Humans communicate non-verbally by sharing physical rhythms, such as nodding and gestures, to involve each other. This sharing of physicality creates a sense of unity and makes humans feel involved with others. In this p…

Motion GenerationUnity

Generalized Modal Analysis in Power System with High CIG Penetration: Concept and Quantitative Assessment

2024-06-24 · Le Zheng, Jiajie Zheng, Chongru Liu

This paper presents a Generalized Modal Analysis (GMA) concept for the small-signal stability analysis of power systems with high penetration of Converter-Interfaced Generation (CIG). GMA quantitatively assesses interact…

Stability-driven Contact Reconstruction From Monocular Color Images

2022-05-02 · CVPR 2022 1 · Zimeng Zhao, Binghui Zuo, Wei Xie, Yangang Wang

Physical contact provides additional constraints for hand-object state reconstruction as well as a basis for further understanding of interaction affordances. Estimating these severely occluded regions from monocular ima…

Object