paper-with-me

홈 › Papers

CamDirector: Towards Long-Term Coherent Video Trajectory Editing

2026-02-27 · Zhihao Shi, Kejia Yin, Weilin Wan, Yuhongze Zhou, Yuanhao Yu, Xinxin Zuo, Qiang Sun, Juwei Lu arxiv

Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regions, upgrading amateur footage into professionally styled videos. Existing VTE methods struggle with precise camera control and long-range consistency because they either inject target poses through a limited-capacity embedding or rely on single-frame warping with only implicit cross-frame aggregation in video diffusion models. To address these issues, we introduce a new VTE framework that 1) explicitly aggregates information across the entire source video via a hybrid warping scheme. Specifically, static regions are progressively fused into a world cache then rendered to target camera poses, while dynamic regions are directly warped; their fusion yields globally consistent coarse frames that guide refinement. 2) processes video segments jointly with their history via a history-guided autoregressive diffusion model, while the world cache is incrementally updated to reinforce already inpainted content, enabling long-term temporal coherence. Finally, we present iPhone-PTZ, a new VTE benchmark with diverse camera motions and large trajectory variations, and achieve state-of-the-art performance with fewer parameters.

📄 PDF Abstract BibTeX arXiv:2603.02256

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models

2025-03-07 · Mark YU, WenBo Hu, Jinbo Xing, Ying Shan

We present TrajectoryCrafter, a novel approach to redirect camera trajectories for monocular videos. By disentangling deterministic view transformations from stochastic content generation, our method achieves precise con…

WorldOlympiad: Can Your World Model Survive a Triathlon?

2026-06-09 · Yuke Zhao, Wangbo Zhao, Weijie Wang, Zeyu Zhang 외 arxiv

We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fidelity. While existing benchmarks often focus on visual quality, sema…

Object Segmentation

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

2026-08-04 · Yibo Yuan, Jiacheng Fu, Jiangtong Zhu, Yi Li 외 arxiv

End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output formats, with limited scalability. Can video generation instead provide …

Scene UnderstandingTrajectory PlanningVideo Generation

Planning with Sketch-Guided Verification for Physics-Aware Video Generation

2025-11-21 · Yidong Huang, Zun Wang, Han Lin, Dong-Ki Kim 외 arxiv

Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-sho…

Video GenerationMotion Planning

Predicting 4D Hand Trajectory from Monocular Videos

2025-01-14 · Yufei Ye, Yao Feng, Omid Taheri, Haiwen Feng 외

We present HaPTIC, an approach that infers coherent 4D hand trajectories from monocular videos. Current video-based hand pose reconstruction methods primarily focus on improving frame-wise 3D pose using adjacent frames r…

Pose Estimation