paper-with-me

홈 › Papers

MAU: A Motion-Aware Unit for Video Prediction and Beyond

2021-12-01 · NeurIPS 2021 12 · Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma, Yan Ye, Xiang Xinguang, Wen Gao

Accurately predicting inter-frame motion information plays a key role in video prediction tasks. In this paper, we propose a Motion-Aware Unit (MAU) to capture reliable inter-frame motion information by broadening the temporal receptive field of the predictive units. The MAU consists of two modules, the attention module and the fusion module. The attention module aims to learn an attention map based on the correlations between the current spatial state and the historical spatial states. Based on the learned attention map, the historical temporal states are aggregated to an augmented motion information (AMI). In this way, the predictive unit can perceive more temporal dynamics from a wider receptive field. Then, the fusion module is utilized to further aggregate the augmented motion information (AMI) and current appearance information (current spatial state) to the final predicted frame. The computation load of MAU is relatively low and the proposed unit can be easily applied to other predictive models. Moreover, an information recalling scheme is employed into the encoders and decoders to help preserve the visual details of the predictions. We evaluate the MAU on both video prediction and early action recognition tasks. Experimental results show that the MAU outperforms the state-of-the-art methods on both tasks.

📄 PDF Abstract BibTeX

Code (1)

ZhengChang467/MAU 공식 구현 pytorch

Tasks

Action RecognitionVideo Prediction

Similar Papers 제목 키워드 기반

STAU: A SpatioTemporal-Aware Unit for Video Prediction and Beyond

2022-04-20 · Zheng Chang, Xinfeng Zhang, Shanshe Wang, Siwei Ma 외

Video prediction aims to predict future frames by modeling the complex spatiotemporal dynamics in videos. However, most of the existing methods only model the temporal information and the spatial information for videos i…

Action Recognitionobject-detectionObject DetectionPrediction+1

ContextVP: Fully Context-Aware Video Prediction

2017-10-23 · ECCV 2018 9 · Wonmin Byeon, Qin Wang, Rupesh Kumar Srivastava, Petros Koumoutsakos

Video prediction models based on convolutional networks, recurrent networks, and their combinations often result in blurry predictions. We identify an important contributing factor for imprecise predictions that has not …

PredictionVideo Prediction

Long History Short-Term Memory for Long-Term Video Prediction

2019-09-25 · Wonmin Byeon, Jan Kautz

While video prediction approaches have advanced considerably in recent years, learning to predict long-term future is challenging — ambiguous future or error propagation over time yield blurry predictions. To address thi…

Video Prediction

HumanNet: Scaling Human-centric Video Learning to One Million Hours

2026-05-07 · Yufan Deng, Daquan Zhou arxiv

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large,…

Representation Learning

Efficient Motion-Aware Video MLLM

2025-01-01 · CVPR 2025 1 · Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo 외

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Mot…

Question AnsweringVideo Question AnsweringVideo Understanding