paper-with-me

Papers

Spatio-temporal video autoencoder with differentiable memory

2015-11-19 · Viorica Patraucean, Ankur Handa, Roberto Cipolla

We describe a new spatio-temporal video autoencoder, based on a classic spatial image autoencoder and a novel nested temporal autoencoder. The temporal encoder is represented by a differentiable visual memory composed of convolutional long short-term memory (LSTM) cells that integrate changes over time. Here we target motion changes and use as temporal decoder a robust optical flow prediction module together with an image sampler serving as built-in feedback loop. The architecture is end-to-end differentiable. At each time step, the system receives as input a video frame, predicts the optical flow based on the current observation and the LSTM memory state as a dense transformation map, and applies it to the current frame to generate the next frame. By minimising the reconstruction error between the predicted next frame and the corresponding ground truth next frame, we train the whole system to extract features useful for motion estimation without any supervision effort. We present one direct application of the proposed framework in weakly-supervised semantic segmentation of videos through label propagation using optical flow.

📄 PDF Abstract BibTeX arXiv:1511.06309

Code (1)

viorik/ConvLSTM 공식 구현

Tasks

DecoderMotion EstimationOptical Flow EstimationSemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic Segmentation

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Solana Customer Service Number +1-833-534-1729 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Pedestrian Spatio-Temporal Information Fusion For Video Anomaly Detection

2022-11-18 · Chao Hu, Liqiang Zhu

Aiming at the problem that the current video anomaly detection cannot fully use the temporal information and ignore the diversity of normal behavior, an anomaly detection method is proposed to integrate the spatiotempora…

Anomaly DetectionDecoderVideo Anomaly Detection

Semi Supervised Meta Learning for Spatiotemporal Learning

2023-07-09 · Faraz Waseem, Pratyush Muthukumar

We approached the goal of applying meta-learning to self-supervised masked autoencoders for spatiotemporal learning in three steps. Broadly, we seek to understand the impact of applying meta-learning to existing state-of…

Action ClassificationClassificationMeta-LearningRepresentation Learning+1

EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens

2022-11-19 · Sunil Hwang, Jaehong Yoon, Youngwan Lee, Sung Ju Hwang

Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memor…

Action RecognitionObject State Change ClassificationObject State Change Classification on Ego4DRepresentation Learning+3

MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion

2024-10-10 · Onkar Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal 외

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video p…

Denoisingparameter-efficient fine-tuningQuantizationText-to-Video Generation+3

S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction

2023-07-13 · Mohammad Adiban, Kalin Stefanov, Sabato Marco Siniscalchi, Giampiero Salvi

We address the video prediction task by putting forth a novel model that combines (i) a novel hierarchical residual learning vector quantized variational autoencoder (HR-VQVAE), and (ii) a novel autoregressive spatiotemp…

PredictionVideo Prediction