3D Human Pose Estimation in Multi-View Operating Room Videos Using Differentiable Camera Projections
3D human pose estimation in multi-view operating room (OR) videos is a relevant asset for person tracking and action recognition. However, the surgical environment makes it challenging to find poses due to sterile clothing, frequent occlusions, and limited public data. Methods specifically designed for the OR are generally based on the fusion of detected poses in multiple camera views. Typically, a 2D pose estimator such as a convolutional neural network (CNN) detects joint locations. Then, the detected joint locations are projected to 3D and fused over all camera views. However, accurate detection in 2D does not guarantee accurate localisation in 3D space. In this work, we propose to directly optimise for localisation in 3D by training 2D CNNs end-to-end based on a 3D loss that is backpropagated through each camera's projection parameters. Using videos from the MVOR dataset, we show that this end-to-end approach outperforms optimisation in 2D space.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Human Pose EstimationAction RecognitionPose EstimationSimilar Papers 제목 키워드 기반
A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
Many approaches have been proposed for human pose estimation in single and multi-view RGB images. However, some environments, such as the operating room, are still very challenging for state-of-the-art RGB methods. In th…
3D Human Pose EstimationMulti-view 3D Human Pose EstimationPose EstimationMVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation
Person detection and pose estimation is a key requirement to develop intelligent context-aware assistance systems. To foster the development of human pose estimation methods and their applications in the Operating Room (…
3D Human Pose Estimation3D Pose EstimationCamera CalibrationHuman Detection+1Next-generation Surgical Navigation: Marker-less Multi-view 6DoF Pose Estimation of Surgical Instruments
State-of-the-art research of traditional computer vision is increasingly leveraged in the surgical domain. A particular focus in computer-assisted surgery is to replace marker-based tracking systems for instrument locali…
AnatomyPose EstimationA Deep Learning Approach for Multi-View Engagement Estimation of Children in a Child-Robot Joint Attention task
In this work we tackle the problem of child engagement estimation while children freely interact with a robot in their room. We propose a deep-based multi-view solution that takes advantage of recent developments in huma…
Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures
Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space…
4k