paper-with-me

Papers

NAIMA: Semantics Aware RGB Guided Depth Super-Resolution

2026-04-06 · Tayyab Nasir, Daochang Liu, Ajmal Mian arxiv

Guided depth super-resolution (GDSR) is a multi-modal approach for depth map super-resolution that relies on a low-resolution depth map and a high-resolution RGB image to restore finer structural details. However, the misleading color and texture cues indicating depth discontinuities in RGB images often lead to artifacts and blurred depth boundaries in the generated depth map. We propose a solution that introduces global contextual semantic priors, generated from pretrained vision transformer token embeddings. Our approach to distilling semantic knowledge from pretrained token embeddings is motivated by their demonstrated effectiveness in related monocular depth estimation tasks. We introduce a Guided Token Attention (GTA) module, which iteratively aligns encoded RGB spatial features with depth encodings, using cross-attention for selectively injecting global semantic context extracted from different layers of a pretrained vision transformer. Additionally, we present an architecture called Neural Attention for Implicit Multi-token Alignment (NAIMA), which integrates DINOv2 with GTA blocks for a semantics-aware GDSR. Our proposed architecture, with its ability to distill semantic knowledge, achieves significant improvements over existing methods across multiple scaling factors and datasets.

📄 PDF Abstract BibTeX arXiv:2604.04407

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth Estimation

Similar Papers 제목 키워드 기반

Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth Estimation

2021-08-19 · ICCV 2021 10 · Hyunyoung Jung, Eunhyeok Park, Sungjoo Yoo

Self-supervised monocular depth estimation has been widely studied, owing to its practical importance and recent promising improvements. However, most works suffer from limited supervision of photometric consistency, esp…

Depth EstimationMetric LearningMonocular Depth Estimation

Learning Depth via Leveraging Semantics: Self-supervised Monocular Depth Estimation with Both Implicit and Explicit Semantic Guidance

2021-02-11 · Rui Li, Xiantuo He, Danna Xue, Shaolin Su 외

Self-supervised depth estimation has made a great success in learning depth from unlabeled image sequences. While the mappings between image and pixel-wise depth are well-studied in current methods, the correlation betwe…

Depth EstimationMonocular Depth Estimation

UniSem: Generalizable Semantic 3D Reconstruction from Sparse Unposed Images

2026-03-18 · Guibiao Liao, Qian Ren, Kaimin Liao, Hua Wang 외 arxiv

Semantic-aware 3D reconstruction from sparse, unposed images remains challenging for feed-forward 3D Gaussian Splatting (3DGS). Existing methods often predict an over-complete set of Gaussian primitives under sparse-view…

3D ReconstructionDepth Estimation

Towards Comprehensive Representation Enhancement in Semantics-guided Self-supervised Monocular Depth Estimation

2022-10-23 · ECCV 2022 10 · Jingyuan Ma, Xiangyu Lei, Nan Liu, Xian Zhao 외

Semantics-guided self-supervised monocular depth estimation has been widely researched, owing to the strong cross-task correlation of depth and semantics. However, since depth estimation and semantic segmentation are fu…

Depth EstimationMetric LearningMonocular Depth EstimationSemantic Segmentation+1

CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion

2025-10-14 · Jinzhou Lin, Jie Zhou, Wenhao Xu, Rongtao Xu 외 arxiv

Semantic Scene Completion (SSC) aims to infer complete 3D geometry and semantics from monocular images, serving as a crucial capability for camera-based perception in autonomous driving. However, existing SSC methods rel…

3D Semantic Scene CompletionAutonomous Driving