paper-with-me

Papers

M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion

2025-05-22 · Nina Shvetsova, Goutam Bhat, Prune Truong, Hilde Kuehne, Federico Tombari

We tackle the problem of monocular-to-stereo video conversion and propose a novel architecture for inpainting and refinement of the warped right view obtained by depth-based reprojection of the input left view. We extend the Stable Video Diffusion (SVD) model to utilize the input left video, the warped right video, and the disocclusion masks as conditioning input to generate a high-quality right camera view. In order to effectively exploit information from neighboring frames for inpainting, we modify the attention layers in SVD to compute full attention for discoccluded pixels. Our model is trained to generate the right view video in an end-to-end manner by minimizing image space losses to ensure high-quality generation. Our approach outperforms previous state-of-the-art methods, obtaining an average rank of 1.43 among the 4 compared methods in a user study, while being 6x faster than the second placed method.

📄 PDF Abstract BibTeX arXiv:2505.16565

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

2024-06-29 · Peng Dai, Feitong Tan, Qiangeng Xu, David Futschik 외

Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free app…

DenoisingVideo GenerationVideo Inpainting

SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-Inpainting

2024-12-16 · Jiale Zhang, Qianxi Jia, Yang Liu, Wei zhang 외

Stereo video conversion aims to transform monocular videos into immersive stereo format. Despite the advancements in novel view synthesis, it still remains two major challenges: i) difficulty of achieving high-fidelity a…

Novel View Synthesis

StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

2024-09-11 · Sijie Zhao, WenBo Hu, Xiaodong Cun, Yong Zhang 외

This paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach over…

Video Inpainting

Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video Generation

2026-07-06 · Jingyi Lu, Kai Han arxiv

Monocular-to-stereo conversion synthesizes stereoscopic content from 2D videos for immersive 3D experiences. In modern Depth-Image-Based Rendering (DIBR) approaches, stereo inpainting of disocclusions is the critical bot…

Self-Supervised LearningVideo Generation

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

2025-08-11 · Peng Dai, Feitong Tan, Qiangeng Xu, Yihua Huang 외 arxiv

While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and trai…

Video GenerationVideo Inpainting