paper-with-me

홈 › Papers

VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing

2025-04-08 · CVPR 2025 1 · Juan Luis Gonzalez Bello, Xu Yao, Alex Whelan, Kyle Olszewski, Hyeongwoo Kim, Pablo Garrido

We present an implicit video representation for occlusions, appearance, and motion disentanglement from monocular videos, which we call Video SPatiotemporal Splines (VideoSPatS). Unlike previous methods that map time and coordinates to deformation and canonical colors, our VideoSPatS maps input coordinates into Spatial and Color Spline deformation fields $D_s$ and $D_c$, which disentangle motion and appearance in videos. With spline-based parametrization, our method naturally generates temporally consistent flow and guarantees long-term temporal consistency, which is crucial for convincing video editing. Using multiple prediction branches, our VideoSPatS model also performs layer separation between the latent video and the selected occluder. By disentangling occlusions, appearance, and motion, our method enables better spatiotemporal modeling and editing of diverse videos, including in-the-wild talking head videos with challenging occlusions, shadows, and specularities while maintaining an appropriate canonical space for editing. We also present general video modeling results on the DAVIS and CoDeF datasets, as well as our own talking head video dataset collected from open-source web videos. Extensive ablations show the combination of $D_s$ and $D_c$ under neural splines can overcome motion and appearance ambiguities, paving the way for more advanced video editing models.

📄 PDF Abstract BibTeX arXiv:2504.07146

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementMotion DisentanglementVideo Editing

Methods 이 논문이 사용한 방법론

Call To Westjet Airlines 설명 없음

Similar Papers 제목 키워드 기반

Continuous-Time Spatiotemporal Calibration of a Rolling Shutter Camera---IMU System

2021-08-16 · Jianzhu Huai, Yuan Zhuang, Qicheng Yuan, Yukai Lin

The rolling shutter (RS) mechanism is widely used by consumer-grade cameras, which are essential parts in smartphones and autonomous vehicles. The RS effect leads to image distortion upon relative motion between a camera…

Autonomous VehiclesVideo Stabilization

VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs

2023-04-12 · ICCV 2023 1 · Moayed Haji Ali, Andrew Bond, Tolga Birdal, Duygu Ceylan 외

We propose $\textbf{VidStyleODE}$, a spatiotemporally continuous disentangled $\textbf{Vid}$eo representation based upon $\textbf{Style}$GAN and Neural-$\textbf{ODE}$s. Effective traversal of the latent space learned by …

Image AnimationVideo EditingVideo Generation

Super-Trajectory for Video Segmentation

2017-02-28 · ICCV 2017 10 · Wenguan Wang, Jianbing Shen, Jianwen Xie, Fatih Porikli

We introduce a novel semi-supervised video segmentation approach based on an efficient video representation, called as "super-trajectory". Each super-trajectory corresponds to a group of compact trajectories that exhibit…

ClusteringSegmentationVideo SegmentationVideo Semantic Segmentation

Compositional Video Understanding with Spatiotemporal Structure-based Transformers

2024-01-01 · CVPR 2024 1 · Hoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol Kim

In this paper we suggest a new novel method to understand complex semantic structures through long video inputs. Conventional methods for understanding videos have been focused on short-term clips and trained to get …

Video Understanding

Revealing Occlusions with 4D Neural Fields

2022-04-22 · CVPR 2022 1 · Basile Van Hoorick, Purva Tendulka, Didac Suris, Dennis Park 외

For computer vision systems to operate in dynamic situations, they need to be able to represent and reason about object permanence. We introduce a framework for learning to estimate 4D visual representations from monocul…

Video Understanding