paper-with-me

홈 › Papers

Any Resolution Any Geometry: From Multi-View To Multi-Patch

2026-03-03 · Wenqing Cui, Zhenyu Li, Mykola Lavreniuk, Jian Shi, Ramzi Idoughi, Xiangjun Tang, Peter Wonka arxiv

Joint estimation of surface normals and depth is essential for holistic 3D scene understanding, yet high-resolution prediction remains difficult due to the trade-off between preserving fine local detail and maintaining global consistency. To address this challenge, we propose the Ultra Resolution Geometry Transformer (URGT), which adapts the Visual Geometry Grounded Transformer (VGGT) into a unified multi-patch transformer for monocular high-resolution depth--normal estimation. A single high-resolution image is partitioned into patches that are augmented with coarse depth and normal priors from pre-trained models, and jointly processed in a single forward pass to predict refined geometric outputs. Global coherence is enforced through cross-patch attention, which enables long-range geometric reasoning and seamless propagation of information across patches within a shared backbone. To further enhance spatial robustness, we introduce a GridMix patch sampling strategy that probabilistically samples grid configurations during training, improving inter-patch consistency and generalization. Our method achieves state-of-the-art results on UnrealStereo4K, jointly improving depth and normal estimation, reducing AbsRel from 0.0582 to 0.0291, RMSE from 2.17 to 1.31, and lowering mean angular error from 23.36 degrees to 18.51 degrees, while producing sharper and more stable geometry. The proposed multi-patch framework also demonstrates strong zero-shot and cross-domain generalization and scales effectively to very high resolutions, offering an efficient and extensible solution for high-quality geometry refinement.

📄 PDF Abstract BibTeX arXiv:2603.03026

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationScene Understanding

Similar Papers 제목 키워드 기반

What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs

2024-01-04 · CVPR 2024 1 · Alex Trevithick, Matthew Chan, Towaki Takikawa, Umar Iqbal 외

3D-aware Generative Adversarial Networks (GANs) have shown remarkable progress in learning to generate multi-view-consistent images and 3D geometries of scenes from collections of 2D images via neural volume rendering. Y…

3D geometryNeural RenderingSuper-Resolution

EpiGRAF: Rethinking training of 3D GANs

2022-06-21 · Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang, Peter Wonka

A very recent trend in generative modeling is building 3D-aware generators from 2D image collections. To induce the 3D bias, such models typically rely on volumetric rendering, which is expensive to employ at high resolu…

3D-Aware Image Synthesis

GAAT: Geometry-Aware Alignment Transformer for Multimodal UAV Perception

2026-08-28 · Jingpu Yang, Debin Tang, Yilin Sun, Fengxian Ji 외 arxiv

Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in opt…

Contrastive LearningScene Understanding

Geometry-Aware Reference Synthesis for Multi-View Image Super-Resolution

2022-07-18 · Ri Cheng, Yuqi Sun, Bo Yan, Weimin Tan 외

Recent multi-view multimedia applications struggle between high-resolution (HR) visual experience and storage or bandwidth constraints. Therefore, this paper proposes a Multi-View Image Super-Resolution (MVISR) task. It …

Image Super-ResolutionSuper-ResolutionVideo Super-Resolution

CoMVS-GS: Collaborative Multi-View Stereo and 3D Gaussian Splatting for Surface Reconstruction

2026-08-19 · Shihan Chen, Junjing Zhang, Qingsong Yan, Haibing Liu 외 arxiv

3D Gaussian Splatting enables efficient novel view synthesis, but accurate mesh reconstruction remains difficult in weakly observed and occluded regions, where Gaussian primitives may grow into unstable or geometrically …

Novel View Synthesis