paper-with-me

Papers Camera Pose Estimation

“Camera Pose Estimation” 태그가 달린 논문 408편 · 필터 해제

GS-CPE: Unified 6-Degree-of-Freedom Camera Pose Estimation via 3D Gaussian Splatting

2026-08-11 · Huaiyuan Weng, Chul Min Yeum, Su-Min Kang arxiv

Despite substantial progress in visual localization, from scene coordinate regression to direct camera pose regression, achieving both robust generalization and high accuracy remain challenging. This study introduces GS-…

Camera Pose EstimationVisual Localization

WAT3R: Feedforward Underwater 3D Reconstruction

2026-07-23 · Jiayi Xu, Jiahao Lu, Ziqiang Zheng, Yihao Tan 외 arxiv

Reliable feedforward underwater 3D reconstruction remains challenging due to severe light attenuation and backscattering, which degrade visual quality and disrupt feature consistency across views, leading to inaccurate m…

Monocular Depth EstimationCamera Pose Estimation3D Reconstruction

OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields

2026-07-12 · Yanqin Jiang, Tengfei Wang, Zhengwei Wang, Chenjie Cao 외 arxiv

Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts t…

Camera Pose EstimationTrajectory PredictionDepth EstimationPoint Tracking

Video Generation Models are General-Purpose Vision Learners

2026-07-10 · Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings 외 arxiv

Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In t…

Text-to-Video GenerationCamera Pose Estimation

NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction

2026-07-08 · Xiangyu Sun, Liu Liu, Seungkwon Yang, Jingbing Han 외 arxiv

Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative cam…

Camera Pose Estimation3D Reconstruction

Vision as Unified Multimodal Generation

2026-07-07 · Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen 외 arxiv

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectu…

Camera Pose Estimationmultimodal generationDepth EstimationImage Generation

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

2026-07-07 · Ruihang Zhang, Felix Taubner, Pooja Ravi, Kiriakos N. Kutulakos 외 arxiv

Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself-…

Camera Pose EstimationPose Tracking

Gen4U: Unifying Video Generation and Understanding via Diffusion

2026-07-07 · Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov 외 arxiv

Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…

Camera Pose EstimationVideo ClassificationDepth EstimationVideo Captioning

SAF3R: Dynamic Sparse Attention for Feed-Forward 3D Reconstruction Transformers

2026-07-03 · Jianing Deng, Yuanzhe Li, Jialu Wang, Song Wang 외 arxiv

Feed-forward 3D reconstruction (F3R) transformers have recently achieved remarkable success. However, scaling them to long image sequences remains challenging, as the quadratic complexity of cross-view global attention q…

Camera Pose Estimation3D Reconstruction

Diversity-aware View Partitioning for Scalable VGGT

2026-07-02 · Jinsoo Park, Donggyu Choi, Ahyun Seo, Minsu cho 외 arxiv

Geometry transformers such as VGGT achieve strong performance by jointly reasoning over multiple views with global attention. However, scaling them to large view collections remains challenging due to the quadratic cost …

Camera Pose Estimationgraph partitioning3D Reconstruction

Multi-THuMBS: Multi-person Tracking of 3D Human Meshes Beyond Video Shots

2026-07-02 · Jeongwan On, Muhammad Salman Ali, Muneeb A. Khan, Sunwoo Park 외 arxiv

Tracking multi-person 3D human meshes from in-the-wild videos is a highly challenging problem due to complex interactions, frequent occlusions, and severe truncation inherent in unconstrained environments. While recent a…

Camera Pose EstimationHuman Mesh Recovery

Planar-SfM: Camera Pose Estimation via Homography Graph Embeddings

2026-06-30 · Gabi Pragier, Matan Karklinsky, David Ungarish, Avi Ben-Cohen arxiv

Structure from Motion (SfM) systems traditionally struggle with planar scenes, where standard epipolar geometry-based methods become degenerate. Rather than viewing planar surfaces as a limitation, we propose a unified f…

Camera Pose Estimation

VOCA: Visual Odometry with Codec Awareness

2026-06-30 · Nouri Alexander Hilscher, Mateo de Mayo, Dominik Muhle, Christoph Otten genannt Hermes 외 arxiv

Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. Nearly all Visual Odometry (VO) and Simultaneous Localization and Map…

Camera Pose EstimationVisual Odometry

Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

2026-06-29 · Xi Li, Linyuan Li, Yan Wu, Tong Rao 외 arxiv

Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synth…

Camera Pose EstimationMulti-Task Learning3D ReconstructionDepth Estimation

G-MASt3R-SfM: Graph-based View Pruning and Multi-stage Optimization for Robust SfM

2026-06-22 · Toshiki Watanabe, Shintaro Ito, Natsuki Takama, Koichi Ito 외 arxiv

Structure from Motion (SfM) is essential for multi-view 3D reconstruction, however, its accuracy heavily relies on the accuracy of image matching. While the recent correspondence matching method, MASt3R, enables robust m…

Multi-View 3D ReconstructionCamera Pose EstimationImage Matching

MoonSplat: Monocular Online Gaussian Splatting with Sim(3) Global Optimization

2026-06-16 · Guo Pu, Yixuan Han, Haofeng Li, Yao Zhang 외 arxiv

Online 3D reconstruction from monocular image sequences is a challenging and ongoing research topic. 3D Gaussian Splatting (3DGS), leveraging its high-quality real-time rendering capability, empowers online 3D reconstruc…

Camera Pose Estimation3D Reconstruction

MVM-IOD: An Industrial Object-Centric Benchmark Dataset for the Evaluation of 3D Reconstruction Methods

2026-06-15 · Robert Langendörfer, Markus Hillemann, Markus Ulrich arxiv

3D object reconstruction, and camera pose estimation in industrial applications are challenging tasks, as errors are costly while the computation time is often limited. The complexity of typical industrial objects furthe…

3D Object ReconstructionCamera Pose Estimation3D ReconstructionPoint Clouds

DarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax

2026-06-09 · Minseong Kweon, Wenyuan Zhao, Nuo Chen, Lulin Liu 외 arxiv

Recent feed-forward 3D reconstruction methods have demonstrated strong performance and flexibility in efficient end-to-end scene geometry estimation from image streams. However, their reliance on visible-light appearance…

Camera Pose Estimation3D Reconstruction

VLM3: Vision Language Models Are Native 3D Learners

2026-05-28 · Zhipeng Cai, Zhuang Liu, Yunyang Xiong, Zechun Liu 외 arxiv

Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still largely relies on exp…

Camera Pose EstimationDepth Estimation

Depth2Pose: A Pose-Based Benchmark for Monocular Depth Estimation without Ground-Truth Depth

2026-05-19 · Viktor Kocur, Sithu Aung, Gabrielle Flood, Yaqing Ding 외 arxiv

Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted depth is increasingly used as an input signal for downstream tasks su…

Monocular Depth EstimationCamera Pose EstimationVisual Localization
1–20 / 408 다음 →