Camera Pose Estimation
1개 벤치마크 · 논문 408편 · 이 태스크의 논문 보기 →
Benchmarks
KITTI Odometry Benchmark
Most implemented
SuperGlue: Learning Feature Matching with Graph Neural Networks
Digging Into Self-Supervised Monocular Depth Estimation
CubeSLAM: Monocular 3D Object SLAM
LoFTR: Detector-Free Local Feature Matching with Transformers
Event-based Stereo Visual Odometry
DGC-Net: Dense Geometric Correspondence Network
Papers
GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting
Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generalization and high accuracy remain challenging. This study introduces GS-…
Camera Pose EstimationVisual LocalizationWAT3R: Feedforward Underwater 3D Reconstruction
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate m…
Monocular Depth EstimationCamera Pose Estimation3D ReconstructionOmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts t…
Camera Pose EstimationTrajectory PredictionDepth EstimationPoint TrackingVideo Generation Models are General-Purpose Vision Learners
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In t…
Text-to-Video GenerationCamera Pose EstimationNoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative cam…
Camera Pose Estimation3D ReconstructionVision as Unified Multimodal Generation
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectu…
Camera Pose Estimationmultimodal generationDepth EstimationImage Generation