paper-with-me

Papers

DINOcular: Self-Supervised Visuospatial Representations

2026-08-27 · Farkhat Almukhamedov, Sami Azirar, Hermann Blum arxiv

We introduce a self-supervised framework for learning joint visuospatial representations from RGB-D observations. While modern vision foundation models are trained almost exclusively on RGB images, many embodied systems have access to explicit depth sensing, which provides geometric information that monocular inputs cannot recover. Our method integrates depth-derived geometric priors with a visual backbone through inter-patch and intra-patch fusion, enabling the model to encode both appearance and spatial structure efficiently. The resulting representation shows promising improvements on 3D awareness while preserving semantic transfer: it outperforms prior methods of comparable scale on multiple 3D geometry benchmarks, and remains competitive when probed for standard RGB-D semantic segmentation tasks.

📄 PDF Abstract BibTeX arXiv:2608.27226

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Segmentation

Similar Papers 제목 키워드 기반

Towards a Human-Centred Cognitive Model of Visuospatial Complexity in Everyday Driving

2020-05-29 · Vasiliki Kondyli, Mehul Bhatt, Jakob Suchan

We develop a human-centred, cognitive model of visuospatial complexity in everyday, naturalistic driving conditions. With a focus on visual perception, the model incorporates quantitative, structural, and dynamic attribu…

Benchmarking

Assessment and treatment of visuospatial neglect using active learning with Gaussian processes regression

2023-09-29 · Ivan De Boi, Elissa Embrechts, Quirine Schatteman, Rudi Penne 외

Visuospatial neglect is a disorder characterised by impaired awareness for visual stimuli located in regions of space and frames of reference. It is often associated with stroke. Patients can struggle with all aspects of…

Active LearningGaussian Processesregression

Towards Visuospatial Cognition via Hierarchical Fusion of Visual Experts

2025-05-18 · Qi Feng

While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, visuospatial cognition - reasoning about spatial layouts, relations, and dynamics - remains a significant challenge. Existing models …

Spatial Reasoning

Visuospatial Cognitive Assistant

2025-05-18 · Qi Feng

Video-based spatial cognition is vital for robotics and embodied AI but challenges current Vision-Language Models (VLMs). This paper makes two key contributions. First, we introduce ViCA (Visuospatial Cognitive Assistant…

Spatial Reasoning

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation

2025-12-02 · Zeqi Xiao, Yiwei Zhao, Lingxiao Li, Yushi Lan 외 arxiv

We investigate whether video generative models can exhibit visuospatial intelligence, a capability central to human cognition, using only visual data. To this end, we present Video4Spatial, a framework showing that video…

Instruction FollowingVideo Generation