paper-with-me

홈 › Papers

VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

2021-05-25 · Findings (ACL) 2022 5 · Ayush Shrivastava, Karthik Gopalakrishnan, Yang Liu, Robinson Piramuthu, Gokhan Tür, Devi Parikh, Dilek Hakkani-Tür

Interactive robots navigating photo-realistic environments need to be trained to effectively leverage and handle the dynamic nature of dialogue in addition to the challenges underlying vision-and-language navigation (VLN). In this paper, we present VISITRON, a multi-modal Transformer-based navigator better suited to the interactive regime inherent to Cooperative Vision-and-Dialog Navigation (CVDN). VISITRON is trained to: i) identify and associate object-level concepts and semantics between the environment and dialogue history, ii) identify when to interact vs. navigate via imitation learning of a binary classification head. We perform extensive pre-training and fine-tuning ablations with VISITRON to gain empirical insights and improve performance on CVDN. VISITRON's ability to identify when to interact leads to a natural generalization of the game-play mode introduced by Roman et al. (arXiv:2005.00728) for enabling the use of such models in different environments. VISITRON is competitive with models on the static CVDN leaderboard and attains state-of-the-art performance on the Success weighted by Path Length (SPL) metric.

📄 PDF Abstract BibTeX arXiv:2105.11589

Code (1)

alexa/visitron 공식 구현 pytorch

Tasks

Binary ClassificationImitation LearningNavigateObjectVision and Language Navigation

Similar Papers 제목 키워드 기반

LERF: Language Embedded Radiance Fields

2023-03-16 · ICCV 2023 1 · Justin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 외

Humans describe the physical world using natural language to refer to specific 3D locations based on a vast range of properties: visual appearance, semantics, abstract associations, or actionable affordances. In this wor…

NeRF

Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions

2025-03-24 · Dong Jing, Nanyi Fei, Zhiwu Lu

In the realm of Large Multi-modal Models (LMMs), the instruction quality during the visual instruction tuning stage significantly influences the performance of modality alignment. In this paper, we assess the instruction…

Sentence

VOILA: An Optimised Dialogue System for Interactively Learning Visually-Grounded Word Meanings (Demonstration System)

2017-08-01 · WS 2017 8 · Yanchao Yu, Arash Eshghi, Oliver Lemon

We present VOILA: an optimised, multi-modal dialogue agent for interactive learning of visually grounded word meanings from a human user. VOILA is: (1) able to learn new visual categories interactively from users from sc…

Active Learning

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

2025-06-25 · Yanzhe Chen, Huasong Zhong, Yan Li, Zhenheng Yang

Unified multimodal large language models (MLLMs) have shown promise in jointly advancing multimodal understanding and generation, with visual codebooks discretizing images into tokens for autoregressive modeling. Existin…

16k

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

2026-06-25 · Xumin Yu, Zuyan Liu, Zhenyu Yang, Yuhao Dong 외 arxiv

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete signals in the same way as text inevitabl…

Representation Learning