paper-with-me

홈 › Papers

Learning Disentangled Representations of Videos with Missing Data

2020-12-01 · NeurIPS 2020 12 · Armand Comas, Chi Zhang, Zlatan Feric, Octavia Camps, Rose Yu

Missing data poses significant challenges while learning representations of video sequences. We present Disentangled Imputed Video autoEncoder (DIVE), a deep generative model that imputes and predicts future video frames in the presence of missing data. Specifically, DIVE introduces a missingness latent variable, disentangles the hidden video representations into static and dynamic appearance, pose, and missingness factors for each object, while it imputes each object trajectory where data is missing. On a moving MNIST dataset with various missing scenarios, DIVE outperforms the state of the art baselines by a substantial margin. We also present comparisons on a real-world MOTSChallenge pedestrian dataset, which demonstrates the practical value of our method in a more realistic setting. Our code can be found in https://github.com/Rose-STL-Lab/DIVE.

📄 PDF Abstract BibTeX

Code (1)

Rose-STL-Lab/DIVE 공식 구현 pytorch

Tasks

Object

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Conditional MoCoGAN for Zero-Shot Video Generation

2021-09-13 · Shun Kimura, Kazuhiko Kawamoto

We propose a conditional generative adversarial network (GAN) model for zero-shot video generation. In this study, we have explored zero-shot conditional generation setting. In other words, we generate unseen videos from…

Generative Adversarial NetworkImage GenerationVideo Generation

Learning Disentangled Representations of Video with Missing Data

2020-06-23 · Armand Comas-Massagué, Chi Zhang, Zlatan Feric, Octavia Camps 외

Missing data poses significant challenges while learning representations of video sequences. We present Disentangled Imputed Video autoEncoder (DIVE), a deep generative model that imputes and predicts future video frames…

A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning

2017-10-16 · NeurIPS 2017 12 · Marco Fraccaro, Simon Kamronn, Ulrich Paquet, Ole Winther

This paper takes a step towards temporal reasoning in a dynamically changing video, not in the pixel space that constitutes its frames, but in a latent space that describes the non-linear dynamics of the objects in its w…

Imputation

Unsupervised Learning of Disentangled Representations from Video

2017-05-31 · NeurIPS 2017 12 · Emily Denton, Vighnesh Birodkar

We present a new model DrNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each f…

DiA-gnostic VLVAE: Disentangled Alignment-Constrained Vision Language Variational AutoEncoder for Robust Radiology Reporting with Missing Modalities

2025-11-08 · Nagur Shareef Shaik, Teja Krishna Cherukuri, Adnan Masood, Dong Hye Ye arxiv

The integration of medical images with clinical context is essential for generating accurate and clinically interpretable radiology reports. However, current automated methods often rely on resource-heavy Large Language …

Knowledge Graphs