paper-with-me

Papers

Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation

2024-07-31 · Sudhir Yarram, Junsong Yuan

Video extrapolation in space and time (VEST) enables viewers to forecast a 3D scene into the future and view it from novel viewpoints. Recent methods propose to learn an entangled representation, aiming to model layered scene geometry, motion forecasting and novel view synthesis together, while assuming simplified affine motion and homography-based warping at each scene layer, leading to inaccurate video extrapolation. Instead of entangled scene representation and rendering, our approach chooses to disentangle scene geometry from scene motion, via lifting the 2D scene to 3D point clouds, which enables high quality rendering of future videos from novel views. To model future 3D scene motion, we propose a disentangled two-stage approach that initially forecasts ego-motion and subsequently the residual motion of dynamic objects (e.g., cars, people). This approach ensures more precise motion predictions by reducing inaccuracies from entanglement of ego-motion with dynamic object motion, where better ego-motion forecasting could significantly enhance the visual outcomes. Extensive experimental analysis on two urban scene datasets demonstrate superior performance of our proposed method in comparison to strong baselines.

📄 PDF Abstract BibTeX arXiv:2407.21450

Code (0)

등록된 구현이 없습니다.

Tasks

Motion ForecastingNovel View Synthesis

Similar Papers 제목 키워드 기반

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

2025-06-16 · Yining Shi, Kun Jiang, Qiang Meng, Ke Wang 외

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (age…

Autonomous DrivingRepresentation Learning

MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds

2024-05-27 · CVPR 2025 1 · Jiahui Lei, Yijia Weng, Adam Harley, Leonidas Guibas 외

We introduce 4D Motion Scaffolds (MoSca), a modern 4D reconstruction system designed to reconstruct and synthesize novel views of dynamic scenes from monocular videos captured casually in the wild. To address such a chal…

4D reconstructionPose Estimation

EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos

2024-05-30 · Masashi Hatano, Ryo Hachiuma, Hideo Saito

Predicting future human behavior from egocentric videos is a challenging but critical task for human intention understanding. Existing methods for forecasting 2D hand positions rely on visual representations and mainly f…

Optical Flow Estimation

The Pose Knows: Video Forecasting by Generating Pose Futures

2017-04-28 · ICCV 2017 10 · Jacob Walker, Kenneth Marino, Abhinav Gupta, Martial Hebert

Current approaches in video forecasting attempt to generate videos directly in pixel space using Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs). However, since these approaches try to model all…

Human Pose ForecastingVideo ForecastingVideo Prediction

VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step

2025-04-02 · CVPR 2025 1 · HanYang Wang, Fangfu Liu, Jiawei Chi, Yueqi Duan

Recovering 3D scenes from sparse views is a challenging task due to its inherent ill-posed problem. Conventional methods have developed specialized solutions (e.g., geometry regularization or feed-forward deterministic m…

DenoisingScene Generation