paper-with-me

Papers

Tempered Self-Similarity Alignment for Physically Plausible Video Generation

2026-05-24 · Manjin Kim, Suha Kwak, Minsu Cho arxiv

Despite remarkable advances in video generative models, they still struggle to generate physically realistic videos, frequently exhibiting appearance drift, implausible motion, and temporal inconsistencies. In this work, we address this limitation by transferring relational knowledge encoded in spatio-temporal self-similarity (STSS) from visual foundation models into video generative models. STSS represents pairwise similarities among features across space and time, revealing the relational structure of how objects interact with other entities throughout a video, effectively capturing real-world dynamics, including object motion and semantic transformations. To transfer this relational knowledge, we propose Tempered Self-similarity Alignment (TSA) loss, which transforms STSS into probabilistic correspondence distributions and trains the video generative model to align its correspondence distributions with those of the visual foundation model on dynamically changing regions. Evaluated on VideoPhy and VideoPhy2 benchmarks, our method demonstrates substantial improvements in physical plausibility across diverse interaction scenarios, validating the effectiveness of transferring relational knowledge for physically realistic video generation.

📄 PDF Abstract BibTeX arXiv:2605.24962

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

2026-02-22 · Wenhao Shen, Hao Wang, Wanqi Yin, Fayao Liu 외 arxiv

Human mesh recovery (HMR) from a single RGB image is inherently ambiguous, as multiple 3D poses can correspond to the same 2D observation. Recent diffusion-based methods tackle this by generating various hypotheses, but …

Human Mesh Recovery

Physically Plausible 3D Human-Scene Reconstruction from Monocular RGB Image using an Adversarial Learning Approach

2023-07-27 · Sandika Biswas, Kejie Li, Biplab Banerjee, Subhasis Chaudhuri 외

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single m…

3D ReconstructionRobot Navigation

Tourbillon: a Physically Plausible Neural Architecture

2021-07-13 · Mohammadamin Tavakoli, Peter Sadowski, Pierre Baldi

In a physical neural system, backpropagation is faced with a number of obstacles including: the need for labeled data, the violation of the locality learning principle, the need for symmetric connections, and the lack of…

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

2026-02-24 · Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu 외 arxiv

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single …

Computational EfficiencyContrastive Learning

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

2026-06-03 · Dingbang Huang, Etienne Vouga, Qixing Huang, Georgios Pavlakos arxiv

In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often…

Reinforcement Learning