To The Point: Correspondence-driven monocular 3D category reconstruction
We present To The Point (TTP), a method for reconstructing 3D objects from a single image using 2D to 3D correspondences learned from weak supervision. We recover a 3D shape from a 2D image by first regressing the 2D positions corresponding to the 3D template vertices and then jointly estimating a rigid camera transform and non-rigid template deformation that optimally explain the 2D positions through the 3D shape projection. By relying on 3D-2D correspondences we use a simple per-sample optimization problem to replace CNN-based regression of camera pose and non-rigid deformation and thereby obtain substantially more accurate 3D reconstructions. We treat this optimization as a differentiable layer and train the whole system in an end-to-end manner. We report systematic quantitative improvements on multiple categories and provide qualitative results comprising diverse shape, pose and texture prediction examples. Project website: https://fkokkinos.github.io/to_the_point/.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction
We introduce ViSER, a method for recovering articulated 3D shapes and dense3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple camera vie…
3D Shape Reconstruction from VideosfCOP: Focal Length Estimation from Category-level Object Priors
In the realm of computer vision, the perception and reconstruction of the 3D world through vision signals heavily rely on camera intrinsic parameters, which have long been a subject of intense research within the communi…
Depth EstimationMonocular Depth EstimationObjectRepresentation LearningCategory-Level 3D Correspondence in Camera Space via Morphable Object Priors
Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estimation, current representations fail to capture the fine-grained sema…
Pose EstimationDynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction
Reconstructing scenes and tracking motion are two sides of the same coin. Tracking points allow for geometric reconstruction [14], while geometric reconstruction of (dynamic) scenes allows for 3D tracking of points over …
Mixed RealityMonocular ReconstructionPoint TrackingRobot NavigationLongDPM: Overlap-Aware 4D Reconstruction from Long Monocular Videos
Recovering a dynamic 3D scene from a long monocular video is crucial for dense geometry, camera motion, and temporal correspondence to remain consistent in a shared coordinate system. Existing methods face two key challe…
Dynamic ReconstructionCamera Pose Estimation