paper-with-me

홈 › Papers

GFlow: Recovering 4D World from Monocular Video

2024-05-28 · Shizun Wang, Xingyi Yang, Qiuhong Shen, Zhenxiang Jiang, Xinchao Wang

Recovering 4D world from monocular video is a crucial yet challenging task. Conventional methods usually rely on the assumptions of multi-view videos, known camera parameters, or static scenes. In this paper, we relax all these constraints and tackle a highly ambitious but practical task: With only one monocular video without camera parameters, we aim to recover the dynamic 3D world alongside the camera poses. To solve this, we introduce GFlow, a new framework that utilizes only 2D priors (depth and optical flow) to lift a video to a 4D scene, as a flow of 3D Gaussians through space and time. GFlow starts by segmenting the video into still and moving parts, then alternates between optimizing camera poses and the dynamics of the 3D Gaussian points. This method ensures consistency among adjacent points and smooth transitions between frames. Since dynamic scenes always continually introduce new visual content, we present prior-driven initialization and pixel-wise densification strategy for Gaussian points to integrate new content. By combining all those techniques, GFlow transcends the boundaries of 4D recovery from causal videos; it naturally enables tracking of points and segmentation of moving objects across frames. Additionally, GFlow estimates the camera poses for each frame, enabling novel view synthesis by changing camera pose. This capability facilitates extensive scene-level or object-level editing, highlighting GFlow's versatility and effectiveness. Visit our project page at: https://littlepure2333.github.io/GFlow

📄 PDF Abstract BibTeX arXiv:2405.18426

Code (0)

등록된 구현이 없습니다.

Tasks

4D reconstructionNovel View SynthesisOptical Flow Estimation

Similar Papers 제목 키워드 기반

World-Coordinate Human Motion Retargeting via SAM 3D Body

2025-12-25 · Zhangzheng Tu, Kailun Su, Shaolong Zhu, Yukun Zheng arxiv

Recovering world-coordinate human motion from monocular videos with humanoid robot retargeting is significant for embodied intelligence and robotics. To avoid complex SLAM pipelines or heavy temporal models, we propose a…

Learning the Depths of Moving People by Watching Frozen People

2019-04-25 · CVPR 2019 6 · Zhengqi Li, Tali Dekel, Forrester Cole, Richard Tucker 외

We present a method for predicting dense depth in scenarios where both a monocular camera and people in the scene are freely moving. Existing methods for recovering depth for dynamic, non-rigid objects from monocular vid…

Depth EstimationDepth Prediction

SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular Video

2025-09-10 · David Stotko, Reinhard Klein arxiv

The reconstruction of three-dimensional dynamic scenes is a well-established yet challenging task within the domain of computer vision. In this paper, we propose a novel approach that combines the domains of 3D geometry …

Physical Simulations3D Reconstruction

FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D Reconstruction

2026-01-26 · Wei Cao, Hao Zhang, Fengrui Tian, Yulun Wu 외 arxiv

Camera redirection aims to replay a dynamic scene from a single monocular video under a user-specified camera trajectory. However, large-angle redirection is inherently ill-posed: a monocular video captures only a narrow…

Visual GroundingVideo GenerationPoint Clouds

Recovering Biomechanical Signals from Missing Keypoints Using Temporal Interpolation in Monocular Gait Analysis

2026-09-09 · Shubham Jariwala arxiv

Monocular pose estimation enables low-cost gait analysis but is sensitive to missing keypoints caused by occlusion, detection errors, or efficiency-driven model reduction. While prior work on recovering missing joints fo…

Pose Estimation