paper-with-me

Papers

Emergence of a Shared Canonical Object Frame from In-the-Wild Videos

2026-06-29 · Tom Fischer, Martin Sundermeyer, Adam Kortylewski, Eddy Ilg arxiv

Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing such frames has traditionally required manual annotation, creating a scaling bottleneck that limits category and instance diversity. We show that a shared canonical frame can instead emerge from self-supervised training on object-centric videos captured in the wild, using only noisy camera poses from Structure-from-Motion. Our key idea is to route all training sequences through a shared geometric bottleneck: a coarse canonical mesh that carries no category-specific detail. By learning dense correspondences from image pixels to this mesh, and estimating per-sequence alignments from noisy SfM geometry, a common canonical frame emerges from multi-view consistency and the semantic priors of the feature extractor, without any canonical pose labels or category conditioning. Trained in a self-supervised manner on 160,000 in-the-wild object videos, our method achieves competitive accuracy on category-level pose estimation benchmarks compared to methods that rely on canonical pose supervision. The code and checkpoint is available on https://github.com/Fischer-Tom/Emergent-Canonical-Frame/.

📄 PDF Abstract BibTeX arXiv:2606.30058

Code (0)

등록된 구현이 없습니다.

Tasks

Pose Estimation

Similar Papers 제목 키워드 기반

3D Congealing: 3D-Aware Image Alignment in the Wild

2024-04-02 · Yunzhi Zhang, Zizhang Li, Amit Raj, Andreas Engelhardt 외

We propose 3D Congealing, a novel problem of 3D-aware alignment for 2D images capturing semantically similar objects. Given a collection of unlabeled Internet images, our goal is to associate the shared semantic parts fr…

Pose Estimation

Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstruction

2020-08-28 · NeurIPS 2020 12 · David Novotny, Roman Shapovalov, Andrea Vedaldi

We propose the Canonical 3D Deformer Map, a new representation of the 3D shape of common object categories that can be learned from a collection of 2D images of independent objects. Our method builds in a novel way on co…

3D ReconstructionObject

WildFusion: Learning 3D-Aware Latent Diffusion Models in View Space

2023-11-22 · Katja Schwarz, Seung Wook Kim, Jun Gao, Sanja Fidler 외

Modern learning-based approaches to 3D-aware image synthesis achieve high photorealism and 3D-consistent viewpoint changes for the generated images. Existing approaches represent instances in a shared canonical space. Ho…

3D-Aware Image Synthesis3D geometryDepth EstimationDepth Prediction+2

3D Surface Reconstruction in the Wild by Deforming Shape Priors from Synthetic Data

2023-02-24 · Nicolai Häni, Jun-Jee Chao, Volkan Isler

Reconstructing the underlying 3D surface of an object from a single image is a challenging problem that has received extensive attention from the computer vision community. Many learning-based approaches tackle this prob…

3D ReconstructionObjectPose EstimationSurface Reconstruction

Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild

2022-10-13 · Kaifeng Zhang, Yang Fu, Shubhankar Borse, Hong Cai 외

While 6D object pose estimation has wide applications across computer vision and robotics, it remains far from being solved due to the lack of annotations. The problem becomes even more challenging when moving to categor…

6D Pose Estimation6D Pose Estimation using RGBObjectPose Estimation+1