paper-with-me

홈 › Papers

Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation

2025-11-27 · Weining Ren, Hongjun Wang, Xiao Tan, Kai Han arxiv

We present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap of all input images to a reference frame coordinate system, along with other auxiliary outputs, in a single forward pass. However, we find that current models struggle with fine geometry and robustness due to (\textit{i}) the scarcity of high-fidelity depth and pose supervision and (\textit{ii}) the inherent geometric misalignment from multi-view pointmap regression. Fin3R jointly tackles two issues with an extra lightweight fine-tuning step. We freeze the decoder, which handles view matching, and fine-tune only the image encoder-the component dedicated to feature extraction. The encoder is enriched with fine geometric details distilled from a strong monocular teacher model on large, unlabeled datasets, using a custom, lightweight LoRA adapter. We validate our method on a wide range of models, including DUSt3R, MASt3R, CUT3R, and VGGT. The fine-tuned models consistently deliver sharper boundaries, recover complex structures, and achieve higher geometric accuracy in both single- and multi-view settings, while adding only the tiny LoRA weights, which leave test-time memory and latency virtually unchanged. Project page: \href{http://visual-ai.github.io/fin3r}{https://visual-ai.github.io/fin3r}

📄 PDF Abstract BibTeX arXiv:2511.22429

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge Distillation3D Reconstruction

Similar Papers 제목 키워드 기반

Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos

2025-07-16 · Kaihua Chen, Tarasha Khurana, Deva Ramanan arxiv

We explore novel-view synthesis for dynamic scenes from monocular videos. Prior approaches rely on costly test-time optimization of 4D representations or do not preserve scene geometry when trained in a feed-forward mann…

Video Inpainting

Sat3R: Satellite DSM Reconstruction via RPC-Aware Depth Fine-tuning

2026-05-08 · Qiaoyi Yang, Chaoyi Zhou, Xi Liu, Run Wang 외 arxiv

Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and large-scale geographic mapping. Existing approaches face a fundamenta…

MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos

2026-07-16 · Ziren Gong, Xiaohan Li, Fabio Tosi, Ninghui Xu 외 arxiv

This paper presents MAGiSt3R, a multi-agent 3D reconstruction framework performing reconstruction and camera tracking for monocular RGB videos at almost 10 FPS. MAGiSt3R relies on a feed-forward model from the 3R family …

3D Reconstruction

Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

2024-12-04 · Hanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba 외

Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse envir…

Novel View Synthesis

DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

2025-06-11 · Chieh Hubert Lin, Zhaoyang Lv, Songyin Wu, Zhen Xu 외

We introduce the Deformable Gaussian Splats Large Reconstruction Model (DGS-LRM), the first feed-forward method predicting deformable 3D Gaussian splats from a monocular posed video of any dynamic scene. Feed-forward sce…

Dynamic Reconstruction