paper-with-me

Papers

TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

2025-03-07 · Mark YU, WenBo Hu, Jinbo Xing, Ying Shan

We present TrajectoryCrafter, a novel approach to redirect camera trajectories for monocular videos. By disentangling deterministic view transformations from stochastic content generation, our method achieves precise control over user-specified camera trajectories. We propose a novel dual-stream conditional video diffusion model that concurrently integrates point cloud renders and source videos as conditions, ensuring accurate view transformations and coherent 4D content generation. Instead of leveraging scarce multi-view videos, we curate a hybrid training dataset combining web-scale monocular videos with static multi-view datasets, by our innovative double-reprojection strategy, significantly fostering robust generalization across diverse scenes. Extensive evaluations on multi-view and large-scale monocular videos demonstrate the superior performance of our method.

📄 PDF Abstract BibTeX arXiv:2503.05638

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis

2025-10-08 · Yihao Zhi, Chenghong Li, Hongjie Liao, Xihe Yang 외 arxiv

Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view synthesizers. Nevertheless, current method…

Monocular Depth EstimationNovel View SynthesisVideo GenerationPoint Clouds

Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

2025-06-16 · Junyoung Seo, Jisang Han, Jaewoo Jung, Siyoon Jin 외

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the…

Novel View Synthesis

Visualizing Skiers' Trajectories in Monocular Videos

2023-04-06 · Matteo Dunnhofer, Luca Sordi, Christian Micheloni

Trajectories are fundamental to winning in alpine skiing. Tools enabling the analysis of such curves can enhance the training activity and enrich broadcasting content. In this paper, we propose SkiTraVis, an algorithm to…

Computational Efficiency

Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera

2024-12-17 · CVPR 2025 1 · Zhengdi Yu, Stefanos Zafeiriou, Tolga Birdal

We propose Dyn-HaMR, to the best of our knowledge, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Reconstructing accurate 3D hand meshes from monocu…

Simultaneous Localization and Mapping

TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

2026-05-02 · Nima Rahmanian, Daniel Kienzle, Thomas Gossard, Dvij Kalaria 외 arxiv

We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides $140+$ hours of reconstructed singles and doubles gameplay from monocular broadcast videos, featuring multimodal annotations like high-quali…