paper-with-me

Papers

3D-RCNN: Instance-Level 3D Object Reconstruction via Render-and-Compare

2018-06-01 · CVPR 2018 6 · Abhijit Kundu, Yin Li, James M. Rehg

We present a fast inverse-graphics framework for instance-level 3D scene understanding. We train a deep convolutional network that learns to map image regions to the full 3D shape and pose of all object instances in the image. Our method produces a compact 3D representation of the scene, which can be readily used for applications like autonomous driving. Many traditional 2D vision outputs, like instance segmentations and depth-maps, can be obtained by simply rendering our output 3D scene model. We exploit class-specific shape priors by learning a low dimensional shape-space from collections of CAD models. We present novel representations of shape and pose, that strive towards better 3D equivariance and generalization. In order to exploit rich supervisory signals in the form of 2D annotations like segmentation, we propose a differentiable Render-and-Compare loss that allows 3D shape and pose to be learned with 2D supervision. We evaluate our method on the challenging real-world datasets of Pascal3D+ and KITTI, where we achieve state-of-the-art results.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object ReconstructionAutonomous DrivingObject ReconstructionScene UnderstandingVehicle Pose Estimation

Similar Papers 제목 키워드 기반

Fusion++: Volumetric Object-Level SLAM

2018-08-25 · John McCormac, Ronald Clark, Michael Bloesch, Andrew J. Davison 외

We propose an online object-level SLAM system which builds a persistent and accurate 3D graph map of arbitrary reconstructed objects. As an RGB-D camera browses a cluttered indoor scene, Mask-RCNN instance segmentations …

Loop Closure DetectionObject

PS-RCNN: Detecting Secondary Human Instances in a Crowd via Primary Object Suppression

2020-03-16 · Zheng Ge, Zequn Jie, Xin Huang, Rong Xu 외

Detecting human bodies in highly crowded scenes is a challenging problem. Two main reasons result in such a problem: 1). weak visual cues of heavily occluded instances can hardly provide sufficient information for accura…

Human DetectionObject Detection

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

2026-06-28 · Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee 외 hf

A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be rec…

Instance SegmentationNovel View Synthesis

MooMIns -- Monocular 3D Reconstruction and Object Pose Estimation from Multiple Instances

2026-06-12 · Robert Langendörfer, Markus Hillemann, Markus Ulrich arxiv

Simultaneous 3D reconstruction and 6D object pose estimation from a single monocular image is an inherently ill-posed problem. In industrial settings, however, multiple instances of an object are often randomly arranged …

Monocular Depth EstimationInstance Segmentation3D ReconstructionPose Estimation

SEMI-PointRend: Improved Semiconductor Wafer Defect Classification and Segmentation as Rendering

2023-02-19 · MinJin Hwang, Bappaditya Dey, Enrique Dehaerne, Sandip Halder 외

In this study, we applied the PointRend (Point-based Rendering) method to semiconductor defect segmentation. PointRend is an iterative segmentation algorithm inspired by image rendering in computer graphics, a new image …

Image SegmentationInstance SegmentationSegmentationSemantic Segmentation