paper-with-me

홈 › Papers

Video Autoencoder: self-supervised disentanglement of static 3D structure and motion

2021-10-06 · ICCV 2021 10 · Zihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong Wang

A video autoencoder is proposed for learning disentan- gled representations of 3D structure and camera pose from videos in a self-supervised manner. Relying on temporal continuity in videos, our work assumes that the 3D scene structure in nearby video frames remains static. Given a sequence of video frames as input, the video autoencoder extracts a disentangled representation of the scene includ- ing: (i) a temporally-consistent deep voxel feature to represent the 3D structure and (ii) a 3D trajectory of camera pose for each frame. These two representations will then be re-entangled for rendering the input video frames. This video autoencoder can be trained directly using a pixel reconstruction loss, without any ground truth 3D or camera pose annotations. The disentangled representation can be applied to a range of tasks, including novel view synthesis, camera pose estimation, and video generation by motion following. We evaluate our method on several large- scale natural video datasets, and show generalization results on out-of-domain images.

📄 PDF Abstract BibTeX arXiv:2110.02951

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationDisentanglementNovel View SynthesisPose EstimationVideo Generation

Similar Papers 제목 키워드 기반

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

2020-05-23 · CVPR 2020 6 · Yizhe Zhu, Martin Renqiang Min, Asim Kadav, Hans Peter Graf

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit the benefits of some readily accessible …

Disentanglement

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

2025-10-07 · Hedi Zisling, Ilan Naiman, Nimrod Berman, Supasorn Suwajanakorn 외 arxiv

Unsupervised representation learning, particularly sequential disentanglement, aims to separate static and dynamic factors of variation in data without relying on labels. This remains a challenging problem, as existing a…

Representation Learning

Multifactor Sequential Disentanglement via Structured Koopman Autoencoders

2023-03-30 · Nimrod Berman, Ilan Naiman, Omri Azencot

Disentangling complex data to its latent factors of variation is a fundamental task in representation learning. Existing work on sequential disentanglement mostly provides two factor representations, i.e., it separates t…

DisentanglementInductive BiasRepresentation Learning

Disentangled Recurrent Wasserstein Autoencoder

2021-01-19 · ICLR 2021 1 · Jun Han, Martin Renqiang Min, Ligong Han, Li Erran Li 외

Learning disentangled representations leads to interpretable models and facilitates data generation with style transfer, which has been extensively studied on static data such as images in an unsupervised learning framew…

DisentanglementRepresentation LearningStyle TransferUnconditional Video Generation+1

Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement Perspective

2022-08-15 · NeurIPS 2023 11 · Pengfei Wei, Lingdong Kong, Xinghua Qu, Yi Ren 외

Unsupervised video domain adaptation is a practical yet challenging task. In this work, for the first time, we tackle it from a disentanglement view. Our key idea is to handle the spatial and temporal domain divergence s…

Action RecognitionDisentanglementDomain Adaptation