Consistent Depth of Moving Objects in Video
We present a method to estimate depth of a dynamic scene, containing arbitrary moving objects, from an ordinary video captured with a moving camera. We seek a geometrically and temporally consistent solution to this underconstrained problem: the depth predictions of corresponding points across frames should induce plausible, smooth motion in 3D. We formulate this objective in a new test-time training framework where a depth-prediction CNN is trained in tandem with an auxiliary scene-flow prediction MLP over the entire input video. By recursively unrolling the scene-flow prediction MLP over varying time steps, we compute both short-range scene flow to impose local smooth motion priors directly in 3D, and long-range scene flow to impose multi-view consistency constraints with wide baselines. We demonstrate accurate and temporally coherent results on a variety of challenging videos containing diverse moving objects (pets, people, cars), as well as camera motion. Our depth maps give rise to a number of depth-and-motion aware video editing effects such as object and lighting insertion.
Code (0)
등록된 구현이 없습니다.
Tasks
Depth EstimationDepth PredictionPredictionRolling Shutter CorrectionVideo EditingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Positional Information is All You Need: A Novel Pipeline for Self-Supervised SVDE from Videos
Recently, much attention has been drawn to learning the underlying 3D structures of a scene from monocular videos in a fully self-supervised fashion. One of the most challenging aspects of this task is handling the indep…
AllDepth EstimationQuantizationUnsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static…
Camera Pose EstimationDepth And Camera MotionDepth EstimationMonocular Depth Estimation+1Coherent Motion Segmentation in Moving Camera Videos using Optical Flow Orientations
In moving camera videos, motion segmentation is commonly performed using the image plane motion of pixels, or optical flow. However, objects that are at different depths from the camera can exhibit different optical flow…
Motion SegmentationOptical Flow EstimationSegmentationEvery Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding
Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network has made significant process recently. Current state-of-the-art (SOTA) methods, are based on the learning fra…
3D geometryDepth And Camera MotionDepth EstimationOptical Flow Estimation+1Decoupling Dynamic Monocular Videos for Dynamic View Synthesis
The challenge of dynamic view synthesis from dynamic monocular videos, i.e., synthesizing novel views for free viewpoints given a monocular video of a dynamic scene captured by a moving camera, mainly lies in accurately …
Optical Flow Estimation