paper-with-me

홈 › Papers

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control

2025-02-14 · Teng Li, Guangcong Zheng, Rui Jiang, Shuigenzhan, Tao Wu, Yehao Lu, Yining Lin, Xi Li

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameters when working with arbitrary real-world images without knowledge of their depth nor scene scale. To address these real-world application issues, we propose RealCam-I2V, a novel diffusion-based video generation framework that integrates monocular metric depth estimation to establish 3D scene reconstruction in a preprocessing step. During training, the reconstructed 3D scene enables scaling camera parameters from relative to absolute values, ensuring compatibility and scale consistency across diverse real-world images. In inference, RealCam-I2V offers an intuitive interface where users can precisely draw camera trajectories by dragging within the 3D scene. To further enhance precise camera control and scene consistency, we propose scene-constrained noise shaping, which shapes high-level noise and also allows the framework to maintain dynamic, coherent video generation in lower noise stages. RealCam-I2V achieves significant improvements in controllability and video quality on the RealEstate10K and out-of-domain images. We further enables applications like camera-controlled looping video generation and generative frame interpolation. We will release our absolute-scale annotation, codes, and all checkpoints. Please see dynamic results in https://zgctroy.github.io/RealCam-I2V.

📄 PDF Abstract BibTeX arXiv:2502.10059

Code (0)

등록된 구현이 없습니다.

Tasks

3D Scene ReconstructionDepth EstimationImage to Video GenerationVideo Generation

Similar Papers 제목 키워드 기반

RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control

2026-05-07 · Youcan Xu, Jiaxin Shi, Zhen Wang, Wensong Song 외 arxiv

Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcasting. However, existing implicit synthes…

Data AugmentationVideo Generation

RealCam-Vid: High-resolution Video Dataset with Dynamic Scenes and Metric-scale Camera Movements

2025-04-11 · Guangcong Zheng, Teng Li, Xianpan Zhou, Xi Li

Recent advances in camera-controllable video generation have been constrained by the reliance on static-scene datasets with relative-scale camera annotations, such as RealEstate10K. While these datasets enable basic view…

Video Generation

RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance

2025-12-28 · Chunyuan Chen, Yunuo Cai, Shujuan Li, Weiyun Liang 외 arxiv

Camouflaged image generation (CIG) has recently emerged as an efficient alternative for acquiring high-quality training data for camouflaged object detection (COD). However, existing CIG methods still suffer from a subst…

Object DetectionImage Generation

An End-to-End Real-World Camera Imaging Pipeline

2024-11-16 · Kepeng Xu, Zijia Ma, Li Xu, Gang He 외

Recent advances in neural camera imaging pipelines have demonstrated notable progress. Nevertheless, the real-world imaging pipeline still faces challenges including the lack of joint optimization in system components, c…

Image CompressionTone Mapping

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

2026-06-17 · Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang 외 arxiv

World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference cos…

Video PredictionVideo GenerationImage Editing