ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild
Estimating the pose of a moving camera from monocular video is a challenging problem, especially due to the presence of moving objects in dynamic environments, where the performance of existing camera pose estimation methods are susceptible to pixels that are not geometrically consistent. To tackle this challenge, we present a robust dense indirect structure-from-motion method for videos that is based on dense correspondence initialized from pairwise optical flow. Our key idea is to optimize long-range video correspondence as dense point trajectories and use it to learn robust estimation of motion segmentation. A novel neural network architecture is proposed for processing irregular point trajectory data. Camera poses are then estimated and optimized with global bundle adjustment over the portion of long-range point trajectories that are classified as static. Experiments on MPI Sintel dataset show that our system produces significantly more accurate camera trajectories compared to existing state-of-the-art methods. In addition, our method is able to retain reasonable accuracy of camera poses on fully static scenes, which consistently outperforms strong state-of-the-art dense correspondence based methods with end-to-end deep learning, demonstrating the potential of dense indirect methods based on optical flow and point trajectories. As the point trajectory representation is general, we further present results and comparisons on in-the-wild monocular videos with complex motion of dynamic objects. Code is available at https://github.com/bytedance/particle-sfm.
Code (1)
Tasks
Camera Pose EstimationMotion SegmentationOptical Flow EstimationPose EstimationSimilar Papers 제목 키워드 기반
DATAP-SfM: Dynamic-Aware Tracking Any Point for Robust Structure from Motion in the Wild
This paper proposes a concise, elegant, and robust pipeline to estimate smooth camera trajectories and obtain dense point clouds for casual videos in the wild. Traditional frameworks, such as ParticleSfM~\cite{zhao2022pa…
Camera Pose EstimationDepth EstimationMotion SegmentationOptical Flow Estimation+2Towards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
RGB-D cameras supply rich and dense visual and spatial information for various robotics tasks such as scene understanding, map reconstruction, and localization. Integrating depth and visual information can aid robots in …
3D Plane Detection3d scene graph generationGraph GenerationPanoptic Segmentation+3ActiveSplat: High-Fidelity Scene Reconstruction through Active Gaussian Splatting
We propose ActiveSplat, an autonomous high-fidelity reconstruction system leveraging Gaussian splatting. Taking advantage of efficient and realistic rendering, the system establishes a unified framework for online mappin…
Active 3D ReconstructionDecision MakingTrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes
We tackle the problem of localizing the traffic surveillance cameras in cooperative perception. To overcome the lack of large-scale real-world intersection datasets, we introduce Carla Intersection, a new simulated datas…
Contrastive LearningImage to Point Cloud RegistrationPoint Cloud RegistrationAgent with Tangent-based Formulation and Anatomical Perception for Standard Plane Localization in 3D Ultrasound
Standard plane (SP) localization is essential in routine clinical ultrasound (US) diagnosis. Compared to 2D US, 3D US can acquire multiple view planes in one scan and provide complete anatomy with the addition of coronal…
AnatomyReinforcement Learning (RL)