Towards Accurate Human Pose Estimation in Videos of Crowded Scenes
Video-based human pose estimation in crowded scenes is a challenging problem due to occlusion, motion blur, scale variation and viewpoint change, etc. Prior approaches always fail to deal with this problem because of (1) lacking of usage of temporal information; (2) lacking of training data in crowded scenes. In this paper, we focus on improving human pose estimation in videos of crowded scenes from the perspectives of exploiting temporal context and collecting new data. In particular, we first follow the top-down strategy to detect persons and perform single-person pose estimation for each frame. Then, we refine the frame-based pose estimation with temporal contexts deriving from the optical-flow. Specifically, for one frame, we forward the historical poses from the previous frames and backward the future poses from the subsequent frames to current frame, leading to stable and accurate human pose estimation in videos. In addition, we mine new data of similar scenes to HIE dataset from the Internet for improving the diversity of training set. In this way, our model achieves best performance on 7 out of 13 videos and 56.33 average w\_AP on test dataset of HIE challenge.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityOptical Flow EstimationPose EstimationSimilar Papers 제목 키워드 기반
JRDB-Pose: A Large-scale Dataset for Multi-Person Pose Estimation and Tracking
Autonomous robotic systems operating in human environments must understand their surroundings to make accurate and safe decisions. In crowded human scenes with close-up human-robot interaction and robot navigation, a dee…
DiversityMulti-Person Pose EstimationMulti-Person Pose Estimation and TrackingPose Estimation+2PromptHMR: Promptable Human Mesh Recovery
Human pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxili…
3D Human Pose EstimationHuman Mesh RecoveryLearning to Estimate Robust 3D Human Mesh from In-the-Wild Crowded Scenes
We consider the problem of recovering a single person's 3D human mesh from in-the-wild crowded scenes. While much progress has been in 3D human mesh estimation, existing methods struggle when test input has crowded scene…
2D Human Pose Estimation3D Human Pose Estimation3D Multi-Person Human Pose Estimation3D Multi-Person Pose Estimation+1Toward Accurate Person-level Action Recognition in Videos of Crowded Scenes
Detecting and recognizing human action in videos with crowded scenes is a challenging problem due to the complex environment and diversity events. Prior works always fail to deal with this problem in two aspects: (1) lac…
Action RecognitionAction Recognition In VideosDiversitySemantic Segmentation+1A Simple Baseline for Pose Tracking in Videos of Crowded Scenes
This paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events\cite{lin2020human}; specifically, here we focus on Track3: Crowd Pose Tracking in Complex Events. Remarkabl…
Multi-Object TrackingObject TrackingOptical Flow EstimationPose Tracking