SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular Input
Stereo video synthesis from a monocular input is a demanding task in the fields of spatial computing and virtual reality. The main challenges of this task lie on the insufficiency of high-quality paired stereo videos for training and the difficulty of maintaining the spatio-temporal consistency between frames. Existing methods primarily address these issues by directly applying novel view synthesis (NVS) techniques to video, while facing limitations such as the inability to effectively represent dynamic scenes and the requirement for large amounts of training data. In this paper, we introduce a novel self-supervised stereo video synthesis paradigm via a video diffusion model, termed SpatialDreamer, which meets the challenges head-on. Firstly, to address the stereo video data insufficiency, we propose a Depth based Video Generation module DVG, which employs a forward-backward rendering mechanism to generate paired videos with geometric and temporal priors. Leveraging data generated by DVG, we propose RefinerNet along with a self-supervised synthetic framework designed to facilitate efficient and dedicated training. More importantly, we devise a consistency control module, which consists of a metric of stereo deviation strength and a Temporal Interaction Learning module TIL for geometric and temporal consistency ensurance respectively. We evaluated the proposed method against various benchmark methods, with the results showcasing its superior performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Novel View SynthesisVideo GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SeLFVi: Self-Supervised Light-Field Video Reconstruction From Stereo Video
Light-field (LF) imaging is appealing to the mobile devices market because of its capability for intuitive post-capture processing. Acquiring LF data with high angular, spatial and temporal resolution poses significa…
Self-Supervised LearningVideo ReconstructionSelf-Supervised Learning of Depth and Ego-motion with Differentiable Bundle Adjustment
Learning to predict scene depth and camera motion from RGB inputs only is a challenging task. Most existing learning based methods deal with this task in a supervised manner which require ground-truth data that is expens…
Depth And Camera MotionSelf-Supervised LearningGeometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation
Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bot…
Self-Supervised LearningVideo GenerationFlow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo Matching
In this paper, we propose a unified method to jointly learn optical flow and stereo matching. Our first intuition is stereo matching can be modeled as a special case of optical flow, and we can leverage 3D geometry behin…
3D geometryOptical Flow EstimationSelf-Supervised LearningStereo MatchingSpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains significantly limited. Current methods o…
Reinforcement LearningScene UnderstandingSpatial Reasoning