On the Role of Rotation Equivariance in Monocular 2D-to-3D Human Pose Lifting
Estimating 3D from 2D is one of the central tasks in computer vision. In this work, we consider the monocular setting, i.e. single-view input, for 3D human pose estimation (HPE), where the goal is to predict a 3D point set of human skeletal joints from a single 2D image, typically via 2D keypoint detection followed by 2D-to-3D lifting. Despite their success, we find that current lifting models exhibit strong performance degradation under rotations. We address this by considering different approaches to incorporating rotation equivariance, including explicit equivariant architectures and standard models. Utilising common HPE benchmarks, we demonstrate that rotation equivariance can be effectively learned via rotation-based data augmentation applied jointly to input and output poses. This significantly improves robustness to rotations and, in this setting, outperforms methods that are fully equivariant by design, while maintaining a lower computational cost.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Human Pose EstimationKeypoint DetectionData AugmentationSimilar Papers 제목 키워드 기반
Scale-Rotation-Equivariant Lie Group Convolution Neural Networks (Lie Group-CNNs)
The weight-sharing mechanism of convolutional kernels ensures translation-equivariance of convolution neural networks (CNNs). Recently, rotation-equivariance has been investigated. However, research on scale-equivariance…
image-classificationImage ClassificationRotated MNISTHarmformer: Harmonic Networks Meet Transformers for Continuous Roto-Translation Equivariance
CNNs exhibit inherent equivariance to image translation, leading to efficient parameter and data usage, faster learning, and improved robustness. The concept of translation equivariant networks has been successfully exte…
TranslationRotationally Equivariant 3D Object Detection
Rotation equivariance has recently become a strongly desired property in the 3D deep learning community. Yet most existing methods focus on equivariance regarding a global input rotation while ignoring the fact that rota…
3D Object DetectionAutonomous DrivingObjectobject-detection+1FRED: Towards a Full Rotation-Equivariance in Aerial Image Object Detection
Rotation-equivariance is an essential yet challenging property in oriented object detection. While general object detectors naturally leverage robustness to spatial shifts due to the translation-equivariance of the conve…
Data AugmentationObjectobject-detectionObject Detection+2On the effectiveness of Rotation-Equivariance in U-Net: A Benchmark for Image Segmentation
Numerous studies have recently focused on incorporating different variations of equivariance in Convolutional Neural Networks (CNNs). In particular, rotation-equivariance has gathered significant attention due to its rel…
Image SegmentationSegmentationSemantic Segmentation