paper-with-me

홈 › Papers

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

2026-03-31 · Ziyang Chen, Yansong Qu, You Shen, Xuan Cheng, Liujuan Cao arxiv

Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research frontier. Contemporary stereo vision backbones typically rely on either Monocular Depth Estimation models or general-purpose Pre-trained Vision Models. Crucially, these models are predominantly pretrained without explicit supervision of camera poses. Given that such geometric knowledge is indispensable for stereo vision, the absence of explicit spatial constraints constitutes a significant performance bottleneck for existing architectures. Recognizing that the Visual Geometry Grounded Transformer (VGGT) operates as a foundation model pre-trained on extensive 3D priors, including camera poses, we investigate its potential as a robust backbone for stereo vision tasks. Nevertheless, empirical results indicate that its direct application to stereo vision yields suboptimal performance. We observe that VGGT suffers from a more significant degradation of geometric details during feature extraction. Such characteristics conflict with the requirements of binocular stereo vision, thereby constraining its efficacy for relative tasks. To bridge this gap, we propose StereoVGGT, a feature backbone specifically tailored for stereo vision. By leveraging the frozen VGGT and introducing a training-free feature adjustment pipeline, we mitigate geometric degradation and harness the latent camera calibration knowledge embedded within the model. StereoVGGT-based stereo matching network achieved the $1^{st}$ rank among all published methods on the KITTI benchmark, validating that StereoVGGT serves as a highly effective backbone for stereo vision.

📄 PDF Abstract BibTeX arXiv:2603.29368

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth Estimation

Similar Papers 제목 키워드 기반

RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

2026-06-16 · Jinhao You, Shuo Lyu, Zhuohang Lyu, Tanxuan Li 외 arxiv

Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators re…

Visual Geometry Transformer in the Wild: Distractor-Free 3D Reconstruction

2026-06-22 · Tianbo Pan, Xingyi Yang, Shizun Wang, Xinchao Wang arxiv

Current end-to-end multi-view 3D reconstruction methods achieve impressive results, but rely on a restrictive static assumption: the scenes is entire distractor-free with perfect cross-view geometry. This reliance on ide…

Multi-View 3D ReconstructionPoint Clouds

FastVGGT: Training-Free Acceleration of Visual Geometry Transformer

2025-09-02 · You Shen, Zhipeng Zhang, Yansong Qu, Xiawu Zheng 외 arxiv

Foundation models for 3D vision have recently demonstrated remarkable capabilities in 3D perception. However, scaling these models to long-sequence image inputs remains a significant challenge due to inference-time ineff…

A Light Touch Approach to Teaching Transformers Multi-view Geometry

2022-11-28 · CVPR 2023 1 · Yash Bhalgat, Joao F. Henriques, Andrew Zisserman

Transformers are powerful visual learners, in large part due to their conspicuous lack of manually-specified priors. This flexibility can be problematic in tasks that involve multiple-view geometry, due to the near-infin…

Retrieval

IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation

2026-05-15 · Yuqi Wu, Tianyu Hu, Wenzhao Zheng, Yuanhui Huang 외 arxiv

Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundation models predict explicit geometry by …

Camera Pose EstimationNovel View Synthesis