Papers Camera Pose Estimation
“Camera Pose Estimation” 태그가 달린 논문 408편 · 필터 해제
GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting
Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generalization and high accuracy remain challenging. This study introduces GS-…
Camera Pose EstimationVisual LocalizationWAT3R: Feedforward Underwater 3D Reconstruction
Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate m…
Monocular Depth EstimationCamera Pose Estimation3D ReconstructionOmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields
Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts t…
Camera Pose EstimationTrajectory PredictionDepth EstimationPoint TrackingVideo Generation Models are General-Purpose Vision Learners
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In t…
Text-to-Video GenerationCamera Pose EstimationNoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction
Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative cam…
Camera Pose Estimation3D ReconstructionVision as Unified Multimodal Generation
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectu…
Camera Pose Estimationmultimodal generationDepth EstimationImage GenerationProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-…
Camera Pose EstimationPose TrackingGen4U: Unifying Video Generation and Understanding via Diffusion
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…
Camera Pose EstimationVideo ClassificationDepth EstimationVideo CaptioningSAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers
Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. However, scaling them to long image sequences remains challenging, as the quadratic complexity of cross-view global attention q…
Camera Pose Estimation3D ReconstructionDiversity-aware View Partitioning for Scalable VGGT
Geometry transformers such as VGGT achieve strong performance by jointly reasoning over multiple views with global attention. However, scaling them to large view collections remains challenging due to the quadratic cost …
Camera Pose Estimationgraph partitioning3D ReconstructionMulti-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots
Tracking multi-person 3D human meshes from in-the-wild videos is a highly challenging problem due to complex interactions, frequent occlusions, and severe truncation inherent in unconstrained environments. While recent a…
Camera Pose EstimationHuman Mesh RecoveryPlanar-SfM: Camera Pose Estimation via Homography Graph Embeddings
Structure from Motion (SfM) systems traditionally struggle with planar scenes, where standard epipolar geometry-based methods become degenerate. Rather than viewing planar surfaces as a limitation, we propose a unified f…
Camera Pose EstimationVOCA: Visual Odometry with Codec Awareness
Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. Nearly all Visual Odometry (VO) and Simultaneous Localization and Map…
Camera Pose EstimationVisual OdometryArgus: Metric Panoramic 3D Reconstruction for Indoor Scenes
Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synth…
Camera Pose EstimationMulti-Task Learning3D ReconstructionDepth EstimationG-MASt3R-SfM: Graph-based View Pruning and Multi-stage Optimization for Robust SfM
Structure from Motion (SfM) is essential for multi-view 3D reconstruction, however, its accuracy heavily relies on the accuracy of image matching. While the recent correspondence matching method, MASt3R, enables robust m…
Multi-View 3D ReconstructionCamera Pose EstimationImage MatchingMoonSplat: Monocular Online Gaussian Splatting with Sim(3) Global Optimization
Online 3D reconstruction from monocular image sequences is a challenging and ongoing research topic. 3D Gaussian Splatting (3DGS), leveraging its high-quality real-time rendering capability, empowers online 3D reconstruc…
Camera Pose Estimation3D ReconstructionMVM-IOD: An Industrial Object-Centric Benchmark Dataset for the Evaluation of 3D Reconstruction Methods
3D object reconstruction, and camera pose estimation in industrial applications are challenging tasks, as errors are costly while the computation time is often limited. The complexity of typical industrial objects furthe…
3D Object ReconstructionCamera Pose Estimation3D ReconstructionPoint CloudsDarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax
Recent feed-forward 3D reconstruction methods have demonstrated strong performance and flexibility in efficient end-to-end scene geometry estimation from image streams. However, their reliance on visible-light appearance…
Camera Pose Estimation3D ReconstructionVLM3: Vision Language Models Are Native 3D Learners
Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on exp…
Camera Pose EstimationDepth EstimationDepth2Pose: A Pose-Based Benchmark for Monocular Depth Estimation without Ground-Truth Depth
Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted depth is increasingly used as an input signal for downstream tasks su…
Monocular Depth EstimationCamera Pose EstimationVisual Localization