paper-with-me

홈 › Papers

RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow

2026-04-21 · Xunpei Sun, Zuoxun Hou, Yi Chang, Gang Chen, Wei-Shi Zheng arxiv

Monocular scene flow estimation aims to recover dense 3D motion from image sequences, yet most existing methods are limited to two-frame inputs, restricting temporal modeling and robustness to occlusions. We propose RAFT-MSF++, a self-supervised multi-frame framework that recurrently fuses temporal features to jointly estimate depth and scene flow. Central to our approach is the Geometry-Motion Feature (GMF), which compactly encodes coupled motion and geometry cues and is iteratively updated for effective temporal reasoning. To ensure the robustness of this temporal fusion against occlusions, we incorporate relative positional attention to inject spatial priors and an occlusion regularization module to propagate reliable motion from visible regions. These components enable the GMF to effectively propagate information even in ambiguous areas. Extensive experiments show that RAFT-MSF++ achieves 24.14% SF-all on the KITTI Scene Flow benchmark, with a 30.99% improvement over the baseline and better robustness in occluded regions. The code is available at https://github.com/sunzunyi/RAFT-MSF-PlusPlus.

📄 PDF Abstract BibTeX arXiv:2604.19349

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Flow Estimation

Similar Papers 제목 키워드 기반

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

2026-05-12 · Jisu Nam, Jahyeok Koo, Soowon Son, Jaewoo Jung 외 arxiv

Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging a…

Scene Understanding3D Reconstruction

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

2026-02-09 · Ruijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han 외 arxiv

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and…

Scene Flow Estimation

MCDS-VSS: Moving Camera Dynamic Scene Video Semantic Segmentation by Filtering with Self-Supervised Geometry and Motion

2024-05-30 · Angel Villar-Corrales, Moritz Austermann, Sven Behnke

Autonomous systems, such as self-driving cars, rely on reliable semantic environment perception for decision making. Despite great advances in video semantic segmentation, existing approaches ignore important inductive b…

Decision MakingScene SegmentationSegmentationSelf-Driving Cars+2

GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors

2025-04-01 · Tian-Xing Xu, Xiangjun Gao, WenBo Hu, Xiaoyu Li 외

Despite remarkable advancements in video depth estimation, existing methods exhibit inherent limitations in achieving geometric fidelity through the affine-invariant predictions, limiting their applicability in reconstru…

4D reconstructionDepth Estimationparameter estimation

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

2025-06-03 · Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark 외

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appe…

3D geometryVideo Generation