Motion Adaptive Pose Estimation From Compressed Videos
Human pose estimation from videos has many real-world applications. Existing methods focus on applying models with a uniform computation profile on fully de- coded frames, ignoring the freely available motion signals and motion-compensation residuals from the compressed stream. A novel model, called Motion Adaptive Pose Net is proposed to exploit the compressed streams to efficiently decode pose sequences from videos. The model incorporates a Motion Compensated ConvLSTM to propagate the spatially aligned features, along with an adaptive gate to dynamically determine if the computationally expensive features should be extracted from fully decoded frames to compensate the motion-warped features, solely based on the residual errors. Leveraging the informative yet readily available signals from compressed streams, we propagate the latent features through our Motion Adaptive Pose Net efficiently. Our model outperforms the state-of-the-art models in pose- estimation accuracy on two widely used datasets with only around half of the computation complexity.
Code (0)
등록된 구현이 없습니다.
Tasks
Motion CompensationPose EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MVFlow: Deep Optical Flow Estimation of Compressed Videos with Motion Vector Prior
In recent years, many deep learning-based methods have been proposed to tackle the problem of optical flow estimation and achieved promising results. However, they hardly consider that most videos are compressed and thus…
Optical Flow EstimationLeveraging Video Coding Knowledge for Deep Video Enhancement
Recent advancements in deep learning techniques have significantly improved the quality of compressed videos. However, previous approaches have not fully exploited the motion characteristics of compressed videos, such as…
Video CompressionVideo EnhancementVideo RestorationCompressed-Domain-Aware Online Video Super-Resolution
In bandwidth-limited online video streaming, videos are usually downsampled and compressed. Although recent online video super-resolution (online VSR) approaches achieve promising results, they are still compute-intensiv…
Video Super-ResolutionYou Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on high…
SentenceTemporal Sentence GroundingA Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution
Online processing of compressed videos to increase their resolutions attracts increasing and broad attention. Video Super-Resolution (VSR) using recurrent neural network architecture is a promising solution due to its ef…
Motion EstimationOptical Flow EstimationSuper-ResolutionVideo Super-Resolution