paper-with-me

홈 › Papers

TexAVi: Generating Stereoscopic VR Video Clips from Text Descriptions

2025-01-02 · Vriksha Srihari, R. Bhavya, Shruti Jayaraman, V. Mary Anita Rajam

While generative models such as text-to-image, large language models and text-to-video have seen significant progress, the extension to text-to-virtual-reality remains largely unexplored, due to a deficit in training data and the complexity of achieving realistic depth and motion in virtual environments. This paper proposes an approach to coalesce existing generative systems to form a stereoscopic virtual reality video from text. Carried out in three main stages, we start with a base text-to-image model that captures context from an input text. We then employ Stable Diffusion on the rudimentary image produced, to generate frames with enhanced realism and overall quality. These frames are processed with depth estimation algorithms to create left-eye and right-eye views, which are stitched side-by-side to create an immersive viewing experience. Such systems would be highly beneficial in virtual reality production, since filming and scene building often require extensive hours of work and post-production effort. We utilize image evaluation techniques, specifically Fr\'echet Inception Distance and CLIP Score, to assess the visual quality of frames produced for the video. These quantitative measures establish the proficiency of the proposed method. Our work highlights the exciting possibilities of using natural language-driven graphics in fields like virtual reality simulations.

📄 PDF Abstract BibTeX arXiv:2501.01156

Code (0)

등록된 구현이 없습니다.

Tasks

Depth Estimation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
BASE 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

T-SVG: Text-Driven Stereoscopic Video Generation

2024-12-12 · Qiao Jin, Xiaodong Chen, Wu Liu, Tao Mei 외

The advent of stereoscopic videos has opened new horizons in multimedia, particularly in extended reality (XR) and virtual reality (VR) applications, where immersive content captivates audiences across various platforms.…

Depth EstimationText-to-Video GenerationVideo GenerationVideo Inpainting

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

2024-06-29 · Peng Dai, Feitong Tan, Qiangeng Xu, David Futschik 외

Video generation models have demonstrated great capabilities of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free app…

DenoisingVideo GenerationVideo Inpainting

Lightweight Multiplane Images Network for Real-Time Stereoscopic Conversion from Planar Video

2024-12-04 · Shanding Diao, Yang Zhao, Yuan Chen, Zhao Zhang 외

With the rapid development of stereoscopic display technologies, especially glasses-free 3D screens, and virtual reality devices, stereoscopic conversion has become an important task to address the lack of high-quality s…

2k

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

2025-08-11 · Peng Dai, Feitong Tan, Qiangeng Xu, Yihua Huang 외 arxiv

While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored challenge. We present a pose-free and trai…

Video GenerationVideo Inpainting

StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

2024-09-11 · Sijie Zhao, WenBo Hu, Xiaodong Cun, Yong Zhang 외

This paper presents a novel framework for converting 2D videos to immersive stereoscopic 3D, addressing the growing demand for 3D content in immersive experience. Leveraging foundation models as priors, our approach over…

Video Inpainting