paper-with-me

홈 › Papers

PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

2026-06-01 · Junlin Long, Zeyu Zhang, Xu Deng, Yiran Wang, Yue Yang, Luke Borgnolo, Maxwell Twelftree, Yang Zhao arxiv

Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household service robotics, assistive robotics, and large-scale autonomous exploration. However, recent attempts to unify vision-and-language navigation (VLN) and object goal navigation (ObjNav) remain at the level of architectural fusion, mixed-task training, and large vision-language pretraining, without examining whether independently trained vision and language encoders may already share a common semantic structure. Moreover, even object-centric topological maps still ground language goals through explicit cross-modal supervision such as CLIP or large vision-language models, leaving open whether such grounding is possible from a purely vision-built map. To address these challenges, we extend the Platonic Representation Hypothesis to embodied navigation and recast vision-only ObjNav, cross-modal ObjNav, and VLN as three different interfaces to the same object-centric semantic manifold. We further introduce PlatonicNav, a training-free framework whose Platonic Topological Map fuses geometric and semantic node distances from a self-supervised visual encoder, and grounds language goals via blind matching without any paired vision-language data. Extensive experiments on simulation benchmarks including HM3D-IIN, OVON, and R2R-CE on MP3D, together with deployment on Unitree Go2, demonstrate that PlatonicNav generalizes across tasks, modalities, and embodiments without explicit cross-modal training. Code: https://github.com/AIGeeksGroup/PlatonicNav. Website: https://aigeeksgroup.github.io/PlatonicNav.

📄 PDF Abstract BibTeX arXiv:2606.01788

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic correspondenceVisual Navigation

Similar Papers 제목 키워드 기반

Agent Journey Beyond RGB: Unveiling Hybrid Semantic-Spatial Environmental Representations for Vision-and-Language Navigation

2024-12-09 · Xuesong Zhang, Yunbo Xu, Jia Li, Zhenzhen Hu 외

Navigating unseen environments based on natural language instructions remains difficult for egocentric agents in Vision-and-Language Navigation (VLN). Existing approaches primarily rely on RGB images for environmental re…

Object LocalizationVision and Language NavigationVision-Language NavigationVisual Navigation

Proof of a perfect platonic representation hypothesis

2025-07-01 · Liu Ziyin, Isaac Chuang arxiv

In this note, we elaborate on and explain in detail the proof given by Ziyin et al. (2025) of the ``perfect" Platonic Representation Hypothesis (PRH) for the embedded deep linear network model (EDLN). We show that if tra…

Representation Learning

Rip-NeRF: Anti-aliasing Radiance Fields with Ripmap-Encoded Platonic Solids

2024-05-03 · Junchen Liu, WenBo Hu, Zhuo Yang, Jianteng Chen 외

Despite significant advancements in Neural Radiance Fields (NeRFs), the renderings may still suffer from aliasing and blurring artifacts, since it remains a fundamental challenge to effectively and efficiently characteri…

NeRF

Platonic Transformers: A Solid Choice For Equivariance

2025-10-03 · Mohammad Mohaiminul Islam, Rishabh Anand, David R. Wessels, Friso de Kruiff 외 arxiv

While widespread, Transformers lack inductive biases for geometric symmetries common in science and computer vision. Existing equivariant methods often sacrifice the efficiency and flexibility that make Transformers so e…

Molecular Property PredictionPoint Clouds

Semantic Match Consistency for Long-Term Visual Localization

2018-09-01 · ECCV 2018 9 · Carl Toft, Erik Stenborg, Lars Hammarstrand, Lucas Brynte 외

Robust and accurate visual localization across large appearance variations due to changes in time of day, seasons, or changes of the environment is a challenging problem which is of importance to application areas such a…

Visual Localization