paper-with-me

홈 › Papers

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

2024-12-12 · CVPR 2025 1 · Linyi Jin, Richard Tucker, Zhengqi Li, David Fouhey, Noah Snavely, Aleksander Holynski

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due to the fundamental difficulty of obtaining ground truth annotations. We present a system for mining high-quality 4D reconstructions from internet stereoscopic, wide-angle videos. Our system fuses and filters the outputs of camera pose estimation, stereo depth estimation, and temporal tracking methods into high-quality dynamic 3D reconstructions. We use this method to generate large-scale data in the form of world-consistent, pseudo-metric 3D point clouds with long-term motion trajectories. We demonstrate the utility of this data by training a variant of DUSt3R to predict structure and 3D motion from real-world image pairs, showing that training on our reconstructed data enables generalization to diverse real-world scenes. Project page: https://stereo4d.github.io

📄 PDF Abstract BibTeX arXiv:2412.09621

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationDepth EstimationPose EstimationStereo Depth Estimation

Similar Papers 제목 키워드 기반

Stereo Vision Based Robot for Remote Monitoring with VR Support

2024-06-27 · Mohamed Fazil M. S., Arockia Selvakumar A., Daniel Schilberg

The machine vision systems have been playing a significant role in visual monitoring systems. With the help of stereovision and machine learning, it will be able to mimic human-like visual system and behaviour towards th…

EgoSampling: Fast-Forward and Stereo for Egocentric Videos

2014-12-11 · CVPR 2015 6 · Yair Poleg, Tavi Halperin, Chetan Arora, Shmuel Peleg

While egocentric cameras like GoPro are gaining popularity, the videos they capture are long, boring, and difficult to watch from start to end. Fast forwarding (i.e. frame sampling) is a natural choice for faster video b…

DynamicStereo: Consistent Dynamic Depth from Stereo Videos

2023-05-03 · CVPR 2023 1 · Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova 외

We consider the problem of reconstructing a dynamic scene observed from a stereo camera. Most existing methods for depth from stereo treat different stereo frames independently, leading to temporally inconsistent depth p…

Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision

2025-12-11 · Wentao Zhou, Xuweiyi Chen, Vignesh Rajagopal, Jeffrey Chen 외 arxiv

The success of foundation models in language and vision motivated research in fully end-to-end robot navigation foundation models (NFMs). NFMs directly map monocular visual input to control actions and ignore mid-level v…

Spatial ReasoningDepth EstimationRobot Navigation

Web Stereo Video Supervision for Depth Prediction from Dynamic Scenes

2019-04-25 · Chaoyang Wang, Simon Lucey, Federico Perazzi, Oliver Wang

We present a fully data-driven method to compute depth from diverse monocular video sequences that contain large amounts of non-rigid objects, e.g., people. In order to learn reconstruction cues for non-rigid scenes, we …

Depth EstimationDepth Prediction