Joint Multiview Segmentation and Localization of RGB-D Images Using Depth-Induced Silhouette Consistency
In this paper, we propose an RGB-D camera localization approach which takes an effective geometry constraint, i.e. silhouette consistency, into consideration. Unlike existing approaches which usually assume the silhouettes are provided, we consider more practical scenarios and generate the silhouettes for multiple views on the fly. To obtain a set of accurate silhouettes, precise camera poses are required to propagate segmentation cues across views. To perform better localization, accurate silhouettes are needed to constrain camera poses. Therefore the two problems are intertwined with each other and require a joint treatment. Facilitated by the available depth, we introduce a simple but effective silhouette consistency energy term that binds traditional appearance-based multiview segmentation cost and RGB-D frame-to-frame matching cost together. Optimization of the problem w.r.t. binary segmentation masks and camera poses naturally fits in the graph cut minimization framework and the Gauss-Newton non-linear least-squares method respectively. Experiments show that the proposed approach achieves state-of-the-arts performance on both tasks of image segmentation and camera localization.
Code (0)
등록된 구현이 없습니다.
Tasks
Camera LocalizationImage SegmentationSegmentationSemantic SegmentationSimilar Papers 제목 키워드 기반
DCHM: Depth-Consistent Human Modeling for Multiview Detection
Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for …
Pedestrian DetectionMultiview DetectionDepth EstimationPoint CloudsMSFormer: A Skeleton-multiview Fusion Method For Tooth Instance Segmentation
Recently, deep learning-based tooth segmentation methods have been limited by the expensive and time-consuming processes of data collection and labeling. Achieving high-precision segmentation with limited datasets is cri…
Contrastive LearningInstance SegmentationSegmentationSemantic SegmentationVarifocal Multiview Images: Capturing and Visual Tasks
Multiview images have flexible field of view (FoV) but inflexible depth of field (DoF). To overcome the limitation of multiview images on visual tasks, in this paper, we present varifocal multiview (VFMV) images with fle…
MVDepthNet: Real-time Multiview Depth Estimation Neural Network
Although deep neural networks have been widely applied to computer vision problems, extending them into multiview depth estimation is non-trivial. In this paper, we present MVDepthNet, a convolutional network to solve th…
Data AugmentationDecoderDepth EstimationHand Keypoint Detection in Single Images using Multiview Bootstrapping
We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an in…
Keypoint Detection