paper-with-me

Papers

ConfCtrl: Enabling Precise Camera Control in Video Diffusion via Confidence-Aware Interpolation

2026-03-10 · Liudi Yang, George Eskandar, Fengyi Shen, Mohammad Altillawi, Yang Bai, Chi Zhang, Ziyuan Liu, Abhinav Valada arxiv

We address the challenge of novel view synthesis from only two input images under large viewpoint changes. Existing regression-based methods lack the capacity to reconstruct unseen regions, while camera-guided diffusion models often deviate from intended trajectories due to noisy point cloud projections or insufficient conditioning from camera poses. To address these issues, we propose ConfCtrl, a confidence-aware video interpolation framework that enables diffusion models to follow prescribed camera poses while completing unseen regions. ConfCtrl initializes the diffusion process by combining a confidence-weighted projected point cloud latent with noise as the conditioning input. It then applies a Kalman-inspired predict-update mechanism, treating the projected point cloud as a noisy measurement and using learned residual corrections to balance pose-driven predictions with noisy geometric observations. This allows the model to rely on reliable projections while down-weighting uncertain regions, yielding stable, geometry-aware generation. Experiments on multiple datasets show that ConfCtrl produces geometrically consistent and visually plausible novel views, effectively reconstructing occluded regions under large viewpoint changes.

📄 PDF Abstract BibTeX arXiv:2603.09819

Code (0)

등록된 구현이 없습니다.

Tasks

Novel View Synthesis

Similar Papers 제목 키워드 기반

CameraCtrl: Enabling Camera Control for Text-to-Video Generation

2024-04-02 · Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein 외

Controllability plays a crucial role in video generation since it allows users to create desired content. However, existing models largely overlooked the precise control of camera pose that serves as a cinematic language…

Text-to-Video GenerationVideo Generation

MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent

2025-02-05 · Xinyao Liao, Xianfang Zeng, Liao Wang, Gang Yu 외

We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fi…

Image to Video GenerationMotion GenerationOptical Flow EstimationVideo Generation

OmniCam: Unified Multimodal Video Generation via Camera Control

2025-04-03 · Xiaoda Yang, Jiayang Xu, Kaixuan Luan, Xinyu Zhan 외

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control ca…

Video Generation

Can video generation replace cinematographers? Research on the cinematic language of generated video

2024-12-16 · Xiaozhe Li, Kai Wu, Siyi Yang, YiZhan Qu 외

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object mo…

Video Generation

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

2026-02-26 · Tianxing Xu, Zixuan Wang, Guangyuan Wang, Li Hu 외 arxiv

World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficulties in two key areas: maintaining long-term content consistency when sce…

3D ReconstructionVideo Generation