paper-with-me

홈 › Papers

Learning Multi-frame and Monocular Prior for Estimating Geometry in Dynamic Scenes

2025-05-03 · Seong Hyeon Park, Jinwoo Shin

In monocular videos that capture dynamic scenes, estimating the 3D geometry of video contents has been a fundamental challenge in computer vision. Specifically, the task is significantly challenged by the object motion, where existing models are limited to predict only partial attributes of the dynamic scenes, such as depth or pointmaps spanning only over a pair of frames. Since these attributes are inherently noisy under multiple frames, test-time global optimizations are often employed to fully recover the geometry, which is liable to failure and incurs heavy inference costs. To address the challenge, we present a new model, coined MMP, to estimate the geometry in a feed-forward manner, which produces a dynamic pointmap representation that evolves over multiple frames. Specifically, based on the recent Siamese architecture, we introduce a new trajectory encoding module to project point-wise dynamics on the representation for each frame, which can provide significantly improved expressiveness for dynamic scenes. In our experiments, we find MMP can achieve state-of-the-art quality in feed-forward pointmap prediction, e.g., 15.1% enhancement in the regression error.

📄 PDF Abstract BibTeX arXiv:2505.01737

Code (0)

등록된 구현이 없습니다.

Tasks

3D geometry

Similar Papers 제목 키워드 기반

MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos

2026-05-28 · Daniel Rho, Jun Myeong Choi, Matthew Thornton, Biswadip Dey 외 arxiv

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, l…

ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy

2025-11-27 · Zhiyi Jiang, Yifu Wang, Xuelian Cheng, Zongyuan Ge arxiv

Estimating 3D geometry from monocular colonoscopy images is challenging due to non-Lambertian surfaces, moving light sources, and large textureless regions. While recent 3D geometric foundation models eliminate the need …

Camera Pose Estimation

Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry

2025-06-16 · Junyoung Seo, Jisang Han, Jaewoo Jung, Siyoon Jin 외

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the…

Novel View Synthesis

Enhanced Scale-aware Depth Estimation for Monocular Endoscopic Scenes with Geometric Modeling

2024-08-14 · Ruofeng Wei, Bin Li, Kai Chen, Yiyao Ma 외

Scale-aware monocular depth estimation poses a significant challenge in computer-aided endoscopic navigation. However, existing depth estimation methods that do not consider the geometric priors struggle to learn the abs…

Depth EstimationMonocular Depth Estimation

BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

2025-08-06 · Tongfan Guan, Jiaxin Guo, Chen Wang, Yun-Hui Liu arxiv

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with amb…

Zero-shot GeneralizationStereo Depth Estimation