paper-with-me

Papers

E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training

2025-12-11 · Qitao Zhao, Hao Tan, Qianqian Wang, Sai Bi, Kai Zhang, Kalyan Sunkavalli, Shubham Tulsiani, Hanwen Jiang arxiv

Self-supervised pre-training has driven rapid progress in foundation models for language, 2D images, and video, yet remains largely unexplored for learning 3D-aware representations from multi-view images. In this paper, we present E-RayZer, a self-supervised 3D vision model that learns geometrically grounded representations directly from unlabeled images. Unlike prior self-supervised methods such as RayZer, which infer 3D indirectly through latent-space view synthesis, E-RayZer operates directly in 3D space, performing self-supervised 3D reconstruction with Explicit geometry. This formulation eliminates shortcut solutions and yields representations that are 3D-aware. To ensure convergence and scalability, we introduce a fine-grained learning curriculum that organizes training from easy to hard samples and harmonizes heterogeneous data sources without any supervision. Experiments show that E-RayZer significantly outperforms RayZer on pose estimation and matches or sometimes surpasses fully supervised reconstruction models such as VGGT. Furthermore, its learned representations outperform leading visual pre-training models (e.g., DINOv3, CroCo v2, VideoMAE V2, and RayZer) on 3D downstream tasks, establishing E-RayZer as a promising paradigm for spatial visual pre-training.

📄 PDF Abstract BibTeX arXiv:2512.10950

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionPose Estimation

Similar Papers 제목 키워드 기반

RayZer: A Self-supervised Large View Synthesis Model

2025-05-01 · Hanwen Jiang, Hao Tan, Peng Wang, Haian Jin 외

We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and unca…

modelNovel View Synthesis

WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments

2026-01-15 · Xuweiyi Chen, Wentao Zhou, Zezhou Cheng arxiv

We present WildRayZer, a self-supervised framework for novel view synthesis (NVS) in dynamic environments where both the camera and objects move. Dynamic content breaks the multi-view consistency that static NVS models r…

Novel View SynthesisPose Estimation

S2F2: Self-Supervised High Fidelity Face Reconstruction from Monocular Image

2022-03-15 · Abdallah Dib, Junghyun Ahn, Cedric Thebault, Philippe-Henri Gosselin 외

We present a novel face reconstruction method capable of reconstructing detailed face geometry, spatially varying face reflectance from a single monocular image. We build our work upon the recent advances of DNN-based au…

3D Face ReconstructionFace ReconstructionSelf-Supervised LearningVocal Bursts Intensity Prediction

Locality-Aware Inter-and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

2022-03-27 · Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang 외

Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pieces, i.e., instance discrimination, loca…

PositionRepresentation LearningVideo Reconstruction

Locality-Aware Inter- and Intra-Video Reconstruction for Self-Supervised Correspondence Learning

2022-01-01 · CVPR 2022 1 · Liulei Li, Tianfei Zhou, Wenguan Wang, Lu Yang 외

Our target is to learn visual correspondence from unlabeled videos. We develop LIIR, a locality-aware inter-and intra-video reconstruction framework that fills in three missing pieces, i.e., instance discrimination, …

PositionRepresentation LearningVideo Reconstruction