paper-with-me

Papers

MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics

2018-08-14 · ECCV 2018 9 · Xinchen Yan, Akash Rastogi, Ruben Villegas, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Ersin Yumer, Honglak Lee

Long-term human motion can be represented as a series of motion modes---motion sequences that capture short-term temporal dynamics---with transitions between them. We leverage this structure and present a novel Motion Transformation Variational Auto-Encoders (MT-VAE) for learning motion sequence generation. Our model jointly learns a feature embedding for motion modes (that the motion sequence can be reconstructed from) and a feature transformation that represents the transition of one motion mode to the next motion mode. Our model is able to generate multiple diverse and plausible motion sequences in the future from the same input. We apply our approach to both facial and full body motion, and demonstrate applications like analogy-based motion transfer and video synthesis.

📄 PDF Abstract BibTeX arXiv:1808.04545

Code (1)

xcyan/eccv18_mtvae tf

Tasks

Human DynamicsHuman Pose Forecastingmotion prediction

Similar Papers 제목 키워드 기반

GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

2024-12-15 · CVPR 2025 1 · Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M B Rezende 외

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object …

Autonomous Driving

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

2025-10-22 · Guowei Xu, Yuxuan Bian, Ailing Zeng, Mingyi Shi 외 arxiv

This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequence-to-sequence manner. OmniMotion-X effic…

Video Representation Learning by Recognizing Temporal Transformations

2020-07-21 · Simon Jenni, Givi Meishvili, Paolo Favaro

We introduce a novel self-supervised learning approach to learn representations of videos that are responsive to changes in the motion dynamics. Our representations can be learned from data without human annotation and p…

Action RecognitionRepresentation LearningSelf-Supervised Learning

IAM: Identity-Aware Human Motion and Shape Joint Generation

2026-04-28 · Wenqi Jia, Zekun Li, Abhay Mittal, Chengcheng Tang 외 arxiv

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and gener…

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction

2026-04-09 · Zhi-Yi Lin, Thomas Markhorst, Jouh Yeong Chew, Xucong Zhang arxiv

Human-like multimodal reaction generation is essential for natural group interactions between humans and embodied AI. However, existing approaches are limited to single-modality or speaking-only responses in dyadic inter…