FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry
We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data design. Stage 1 learns a high-fidelity, affine-invariant geometry model by initializing with DINOv3 and training on a curated 10.2M-sample multi-domain corpus with complementary local-detail supervision, yielding sharp boundaries and strong cross-domain generalization. Stage 2 moves beyond global scaling by introducing lightweight pixel-wise calibration fields for metric estimation: a scale field for spatially varying metric alignment and a ray-direction correction field that mitigates directional bias in point-map geometry, together producing metrically consistent 3D point maps. Beyond model design, we identify camera intrinsic coverage, especially focal length distribution mismatch between training and test data, as a key bottleneck for zero-shot metric generalization: performance drops sharply when test intrinsics fall outside the training distribution. To address this, we synthesize additional training data across diverse focal lengths using a Blender-based data engine, repairing under-covered focal regimes and improving robustness under intrinsic shift. Extensive zero-shot evaluations across seven benchmarks show that FoundationGeo significantly strengthens cross-domain robustness, staying near the top across diverse domains while avoiding the sharp cross-domain performance drops observed in other methods. This consistency translates into the best overall performance, surpassing heavier baselines by over 5.2% on average.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationSimilar Papers 제목 키워드 기반
Pixel-wise Attentional Gating for Parsimonious Pixel Labeling
To achieve parsimonious inference in per-pixel labeling tasks with a limited computational budget, we propose a \emph{Pixel-wise Attentional Gating} unit (\emph{PAG}) that learns to selectively process a subset of spatia…
Boundary DetectionSemantic SegmentationSurface Normal EstimationMonocular Piecewise Depth Estimation in Dynamic Scenes by Exploiting Superpixel Relations
In this paper, we propose a novel and specially designed method for piecewise dense monocular depth estimation in dynamic scenes. We utilize spatial relations between neighboring superpixels to solve the inherent relativ…
Depth EstimationMonocular Depth EstimationOptical Flow EstimationSemantic Segmentation+1Fieldscale: Locality-Aware Field-based Adaptive Rescaling for Thermal Infrared Image
Thermal infrared (TIR) cameras are emerging as promising sensors in safety-related fields due to their robustness against external illumination. However, RAW TIR image has 14 bits of pixel depth and needs to be rescaled …
Image Quality AssessmentA Self-Supervised Method for Attenuating Seismic Random and Tracewise Coherent Noise under the Non-Pixelwise Independence Assumption
The attenuation of seismic field noise using self-supervised deep learning has gained attention due to its label-free training process. However, common self-supervised methods are limited by the pixelwise independence as…
DenoisingGeophysicsA Self-Supervised Method for Attenuating Seismic Random and Tracewise Coherent Noise under the Non-Pixelwise Independence Assumption
The attenuation of seismic field noise using self-supervised deep learning has gained attention due to its label-free training process. However, common self-supervised methods are limited by the pixelwise independence as…