Self-supervised Learning of 3D Objects from Natural Images
We present a method to learn single-view reconstruction of the 3D shape, pose, and texture of objects from categorized natural images in a self-supervised manner. Since this is a severely ill-posed problem, carefully designing a training method and introducing constraints are essential. To avoid the difficulty of training all elements at the same time, we propose training category-specific base shapes with fixed pose distribution and simple textures first, and subsequently training poses and textures using the obtained shapes. Another difficulty is that shapes and backgrounds sometimes become excessively complicated to mistakenly reconstruct textures on object surfaces. To suppress it, we propose using strong regularization and constraints on object surfaces and background images. With these two techniques, we demonstrate that we can use natural image collections such as CIFAR-10 and PASCAL objects for training, which indicates the possibility to realize 3D object reconstruction on diverse object categories beyond synthetic datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object ReconstructionObjectObject ReconstructionSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Affinity-based Attention in Self-supervised Transformers Predicts Dynamics of Object Grouping in Humans
The spreading of attention has been proposed as a mechanism for how humans group features to segment objects. However, such a mechanism has not yet been implemented and tested in naturalistic images. Here, we leverage th…
ObjectRepresentation LearningTime to augment self-supervised visual representation learning
Biological vision systems are unparalleled in their ability to learn visual representations without supervision. In machine learning, self-supervised learning (SSL) has led to major advances in forming object representat…
Contrastive LearningObjectRepresentation LearningSelf-Supervised LearningUnsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
The recent enthusiasm for open-world vision systems show the high interest of the community to perform perception tasks outside of the closed-vocabulary benchmark setups which have been so popular until now. Being able t…
ObjectObject LocalizationUnsupervised Object LocalizationPooDLe: Pooled and dense self-supervised learning from naturalistic videos
Self-supervised learning has driven significant progress in learning from single-subject, iconic images. However, there are still unanswered questions about the use of minimally-curated, naturalistic video data, which co…
Optical Flow EstimationSelf-Supervised LearningObject-centric LeJEPA
Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets. Aligning representations at the level o…