Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views
Object viewpoint estimation from 2D images is an essential task in computer vision. However, two issues hinder its progress: scarcity of training data with viewpoint annotations, and a lack of powerful features. Inspired by the growing availability of 3D models, we propose a framework to address both issues by combining render-based image synthesis and CNNs. We believe that 3D models have the potential in generating a large number of images of high variation, which can be well exploited by deep CNN with a high learning capacity. Towards this goal, we propose a scalable and overfit-resistant image synthesis pipeline, together with a novel CNN specifically tailored for the viewpoint estimation task. Experimentally, we show that the viewpoint estimation from our pipeline can significantly outperform state-of-the-art methods on PASCAL 3D+ benchmark.
Code (4)
Tasks
Image GenerationViewpoint EstimationSimilar Papers 제목 키워드 기반
Understanding deep features with computer-generated imagery
We introduce an approach for analyzing the variation of features generated by convolutional neural networks (CNNs) with respect to scene factors that occur in natural images. Such factors may include object style, 3D vie…
Aperture Supervision for Monocular Depth Estimation
We present a novel method to train machine learning algorithms to estimate scene depths from a single image, by using the information provided by a camera's aperture as supervision. Prior works use a depth sensor's outpu…
Depth EstimationMonocular Depth EstimationGGS: Generalizable Gaussian Splatting for Lane Switching in Autonomous Driving
We propose GGS, a Generalizable Gaussian Splatting method for Autonomous Driving which can achieve realistic rendering under large viewpoint changes. Previous generalizable 3D gaussian splatting methods are limited to re…
Autonomous DrivingDepth EstimationCapsules as viewpoint learners for human pose estimation
The task of human pose estimation (HPE) deals with the ill-posed problem of estimating the 3D position of human joints directly from images and videos. In recent literature, most of the works tackle the problem mostly by…
Multi-class ClassificationPose EstimationHow useful is photo-realistic rendering for visual learning?
Data seems cheap to get, and in many ways it is, but the process of creating a high quality labeled dataset from a mass of data is time-consuming and expensive. With the advent of rich 3D repositories, photo-realistic …
Domain AdaptationViewpoint Estimation