paper-with-me

홈 › Papers

LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation

2026-02-07 · Nitesh Subedi, Adam Haroon, Samuel Tetteh, Prajwal Koirala, Cody Fleming, Soumik Sarkar arxiv

We propose LCLA (Language-Conditioned Latent Alignment), a framework for vision-language navigation that learns modular perception-action interfaces by aligning sensory observations to a latent representation of an expert policy. The expert is first trained with privileged state information, inducing a latent space sufficient for control, after which its latent interface and action head are frozen. A lightweight adapter is then trained to map raw visual-language observations, via a frozen vision-language model, into the expert's latent space, reducing the problem of visuomotor learning to supervised latent alignment rather than end-to-end policy optimization. This decoupling enforces a stable contract between perception and control, enabling expert behavior to be reused across sensing modalities and environmental variations. We instantiate LCLA and evaluate it on a vision-language indoor navigation task, where aligned latent spaces yield strong in-distribution performance and robust zero-shot generalization to unseen environments, lighting conditions, and viewpoints while remaining lightweight at inference time.

📄 PDF Abstract BibTeX arXiv:2602.07629

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationZero-shot Generalization

Similar Papers 제목 키워드 기반

SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning

2026-06-11 · Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su 외 arxiv

Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attempt to address this by…

Spatial Reasoning

AlcLaM: Arabic Dialectal Language Model

2024-07-18 · Murtadha Ahmed, Saghir Alfasly, Bo Wen, Jamaal Qasem 외

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inferen…

Language ModelingLanguage Modellingmodel

VisualClaw: A Real-Time, Personalized Agent for the Physical World

2026-06-15 · Haoqin Tu, Jianwen Chen, Zijun Wang, Siwei Han 외 arxiv

Vision language models are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cost when processing dense video frames and…

Open-world Multi-label Text Classification with Extremely Weak Supervision

2024-07-08 · Xintong Li, Jinya Jiang, Ria Dharmani, Jayanth Srinivasa 외

We study open-world multi-label text classification under extremely weak supervision (XWS), where the user only provides a brief description for classification objectives without any labels or ground-truth label space. S…

Keyword ExtractionLanguage ModellingLarge Language ModelMulti-Label Classification+5

AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference

2026-03-23 · Risa Shinoda, Kaede Shiohara, Nakamasa Inoue, Hiroaki Santo 외 arxiv

Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring. Recent advances in deep learning have …