paper-with-me

홈 › Papers

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

2026-07-30 · Jin Cao, Zian Meng, Kaipeng Zhang arxiv

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io

📄 PDF Abstract BibTeX arXiv:2607.28362

Code (3)

AlayaLab/ShadowDancer ★ 17
BaiShuanghao/my_arXiv_daily ★ 205
Tavish9/awesome-daily-AI-arxiv ★ 112

Similar Papers 제목 키워드 기반

Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning

2025-01-13 · Juntao Ren, Priya Sundaresan, Dorsa Sadigh, Sanjiban Choudhury 외

Teaching robots to autonomously complete everyday tasks remains a challenge. Imitation Learning (IL) is a powerful approach that imbues robots with skills via demonstrations, but is limited by the labor-intensive process…

Few-Shot Imitation LearningImitation Learning

MUTLA: A Large-Scale Dataset for Multimodal Teaching and Learning Analytics

2019-10-05 · Fangli Xu, Lingfei Wu, KP Thai, Carol Hsu 외

Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching an…

EEGElectroencephalogram (EEG)

CLUE: Contextualised Unified Explainable Learning of User Engagement in Video Lectures

2022-01-14 · Sujit Roy, Gnaneswara Rao Gorle, Vishal Gaur, Haider Raza 외

Predicting contextualised engagement in videos is a long-standing problem that has been popularly attempted by exploiting the number of views or the associated likes using different computational methods. The recent deca…

Transfer Learning

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

2026-01-09 · Nate Gillman, Yinghua Zhou, Zitian Tang, Evan Luo 외 arxiv

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a cha…

Zero-shot GeneralizationVideo Generation

From Motion Signals to Insights: A Unified Framework for Student Behavior Analysis and Feedback in Physical Education Classes

2025-03-09 · Xian Gao, Jiacheng Ruan, Jingsheng Gao, Mingye Xie 외

Analyzing student behavior in educational scenarios is crucial for enhancing teaching quality and student engagement. Existing AI-based models often rely on classroom video footage to identify and analyze student behavio…

Activity RecognitionHuman Activity Recognition