paper-with-me

Papers

Self-Supervised Decomposition, Disentanglement and Prediction of Video Sequences while Interpreting Dynamics: A Koopman Perspective

2021-10-01 · Armand Comas, Sandesh Ghimire, Haolin Li, Mario Sznaier, Octavia Camps

Human interpretation of the world encompasses the use of symbols to categorize sensory inputs and compose them in a hierarchical manner. One of the long-term objectives of Computer Vision and Artificial Intelligence is to endow machines with the capacity of structuring and interpreting the world as we do. Towards this goal, recent methods have successfully been able to decompose and disentangle video sequences into their composing objects and dynamics, in a self-supervised fashion. However, there has been a scarce effort in giving interpretation to the dynamics of the scene. We propose a method to decompose a video into moving objects and their attributes, and model each object's dynamics with linear system identification tools, by means of a Koopman embedding. This allows interpretation, manipulation and extrapolation of the dynamics of the different objects by employing the Koopman operator K. We test our method in various synthetic datasets and successfully forecast challenging trajectories while interpreting them.

📄 PDF Abstract BibTeX arXiv:2110.00547

Code (0)

등록된 구현이 없습니다.

Tasks

Disentanglement

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled Videos

2019-10-01 · ICCV 2019 10 · Fanyi Xiao, Haotian Liu, Yong Jae Lee

We propose a novel approach that disentangles the identity and pose of objects for image generation. Our model takes as input an ID image and a pose image, and generates an output image with the identity of the ID image …

DisentanglementDiversityImage Generation

Learning to Decompose and Disentangle Representations for Video Prediction

2018-06-11 · NeurIPS 2018 12 · Jun-Ting Hsieh, Bingbin Liu, De-An Huang, Li Fei-Fei 외

Our goal is to predict future video frames given a sequence of input frames. Despite large amounts of video data, this remains a challenging task because of the high-dimensionality of video frames. We address this challe…

DisentanglementPredict Future Video FramesVideo Prediction

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

2020-05-23 · CVPR 2020 6 · Yizhe Zhu, Martin Renqiang Min, Asim Kadav, Hans Peter Graf

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit the benefits of some readily accessible …

Disentanglement

Video Reenactment as Inductive Bias for Content-Motion Disentanglement

2021-01-30 · Juan F. Hernández Albarracín, Adín Ramírez Rivera

Independent components within low-dimensional representations are essential inputs in several downstream tasks, and provide explanations over the observed data. Video-based disentangled factors of variation provide low-d…

DisentanglementInductive BiasMotion Disentanglement

Sample and Predict Your Latent: Modality-free Sequential Disentanglement via Contrastive Estimation

2023-05-25 · Ilan Naiman, Nimrod Berman, Omri Azencot

Unsupervised disentanglement is a long-standing challenge in representation learning. Recently, self-supervised techniques achieved impressive results in the sequential setting, where data is time-dependent. However, the…

DisentanglementRepresentation LearningTime Series