Light3R-SfM: Towards Feed-forward Structure-from-Motion
We present Light3R-SfM, a feed-forward, end-to-end learnable framework for efficient large-scale Structure-from-Motion (SfM) from unconstrained image collections. Unlike existing SfM solutions that rely on costly matching and global optimization to achieve accurate 3D reconstructions, Light3R-SfM addresses this limitation through a novel latent global alignment module. This module replaces traditional global optimization with a learnable attention mechanism, effectively capturing multi-view constraints across images for robust and precise camera pose estimation. Light3R-SfM constructs a sparse scene graph via retrieval-score-guided shortest path tree to dramatically reduce memory usage and computational overhead compared to the naive approach. Extensive experiments demonstrate that Light3R-SfM achieves competitive accuracy while significantly reducing runtime, making it ideal for 3D reconstruction tasks in real-world applications with a runtime constraint. This work pioneers a data-driven, feed-forward SfM approach, paving the way toward scalable, accurate, and efficient 3D reconstruction in the wild.
Code (0)
등록된 구현이 없습니다.
Tasks
3D ReconstructionCamera Pose Estimationglobal-optimizationPose EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Global Structure-from-Motion Meets Feedforward Reconstruction
Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a central challenge in computer vision, with many open problems yet to be solve…
3D ReconstructionA Kernel-Based Identification Approach to LPV Feedforward: With Application to Motion Systems
The increasing demands for motion control result in a situation where Linear Parameter-Varying (LPV) dynamics have to be taken into account. Inverse-model feedforward control for LPV motion systems is challenging, since …
SchedulingGGPT: Geometry Grounded Point Transformer
Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limit…
3D ReconstructionPoint CloudsOmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts t…
Camera Pose EstimationTrajectory PredictionDepth EstimationPoint TrackingEmergence of robust looming selectivity via coordinated inhibitory neural computations
In the locust's lobula giant movement detector neural pathways, four categories of inhibition, i.e., global inhibition, self-inhibition, lateral inhibition, and feed-forward inhibition, have been functionally explored in…