Unsupervised learning of object landmarks by factorized spatial embeddings
Learning automatically the structure of object categories remains an important open problem in computer vision. In this paper, we propose a novel unsupervised approach that can discover and learn landmarks in object categories, thus characterizing their structure. Our approach is based on factorizing image deformations, as induced by a viewpoint change or an object deformation, by learning a deep neural network that detects landmarks consistently with such visual effects. Furthermore, we show that the learned landmarks establish meaningful correspondences between different object instances in a category without having to impose this requirement explicitly. We assess the method qualitatively on a variety of object types, natural and man-made. We also show that our unsupervised landmarks are highly predictive of manually-annotated landmarks in face benchmark datasets, and can be used to regress these with a high degree of accuracy.
Code (1)
Tasks
ObjectUnsupervised Facial Landmark DetectionUnsupervised Human Pose EstimationUnsupervised KeypointsSimilar Papers 제목 키워드 기반
Unsupervised Disentanglement of Pose, Appearance and Background from Images and Videos
Unsupervised landmark learning is the task of learning semantic keypoint-like representations without the use of expensive input keypoint-level annotations. A popular approach is to factorize an image into a pose and app…
DisentanglementVideo PredictionUnsupervised Learning of Object Landmarks via Self-Training Correspondence
This paper addresses the problem of unsupervised discovery of object landmarks. We take a different path compared to that of existing works, based on 2 novel perspectives: (1) Self-training: starting from generic keypoin…
ClusteringObjectUnsupervised Landmark DetectionUnsupervised Discovery of Facial Landmarks and Head Pose
Unsupervised landmark and head pose estimation is fundamental in fields like biometrics, augmented reality, and emotion recognition, offering accurate spatial data without relying on labeled datasets. It enhances sca…
Emotion RecognitionHead Pose EstimationLandmark TrackingPose Estimation+1SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and Decomposition
The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation l…
ObjectRepresentation LearningUnsupervised Discovery of Object Landmarks as Structural Representations
Deep neural networks can model images with rich latent representations, but they cannot naturally conceptualize structures of object categories in a human-perceptible way. This paper addresses the problem of learning obj…
ObjectUnsupervised Facial Landmark DetectionUnsupervised Human Pose EstimationUnsupervised Keypoint Estimation