paper-with-me

홈 › Papers

NOLO: Navigate Only Look Once

2024-08-02 · Bohan Zhou, Zhongbin Zhang, Jiangxing Wang, Zongqing Lu

The in-context learning ability of Transformer models has brought new possibilities to visual navigation. In this paper, we focus on the video navigation setting, where an in-context navigation policy needs to be learned purely from videos in an offline manner, without access to the actual environment. For this setting, we propose Navigate Only Look Once (NOLO), a method for learning a navigation policy that possesses the in-context ability and adapts to new scenes by taking corresponding context videos as input without finetuning or re-training. To enable learning from videos, we first propose a pseudo action labeling procedure using optical flow to recover the action label from egocentric videos. Then, offline reinforcement learning is applied to learn the navigation policy. Through extensive experiments on different scenes both in simulation and the real world, we show that our algorithm outperforms baselines by a large margin, which demonstrates the in-context learning ability of the learned policy. For videos and more information, visit https://sites.google.com/view/nol0.

📄 PDF Abstract BibTeX arXiv:2408.01384

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningNavigateOptical Flow EstimationVisual Navigation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Multi-Head Attention 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Local Navigation and Docking of an Autonomous Robot Mower using Reinforcement Learning and Computer Vision

2021-01-15 · Ali Taghibakhshi, Nathan Ogden, Matthew West

We demonstrate a successful navigation and docking control system for the John Deere Tango autonomous mower, using only a single camera as the input. This vision-only system is of interest because it is inexpensive, simp…

Navigateobject-detectionObject DetectionReinforcement Learning (RL)

Gaze-contingent decoding of human navigation intention on an autonomous wheelchair platform

2021-03-04 · Mahendran Subramanian, Suhyung Park, Pavel Orlov, Ali Shafti 외

We have pioneered the Where-You-Look-Is Where-You-Go approach to controlling mobility platforms by decoding how the user looks at the environment to understand where they want to navigate their mobility device. However, …

Motor ImageryNavigateObject

Navigation of micro-robot swarms for targeted delivery using reinforcement learning

2023-06-30 · Akshatha Jagadish, Manoj Varma

Micro robotics is quickly emerging to be a promising technological solution to many medical treatments with focus on targeted drug delivery. They are effective when working in swarms whose individual control is mostly in…

Navigatereinforcement-learningReinforcement Learning (RL)

Instruction Tuning Chronologically Consistent Language Models

2025-10-13 · Songrun He, Linying Lv, Asaf Manela, Jimmy Wu arxiv

We introduce a family of chronologically consistent, instruction-tuned large language models to eliminate lookahead bias. Each model is trained only on data available before a clearly defined knowledge-cutoff date, ensur…

YOLO11 to Its Genesis: A Decadal and Comprehensive Review of The You Only Look Once (YOLO) Series

2024-06-12 · Ranjan Sapkota, Rizwan Qureshi, Marco Flores Calero, Chetan Badjugar 외

Given the rapid emergence and applications of Large Language This review systematically examines the progression of the You Only Look Once (YOLO) object detection algorithms from YOLOv1 to the recently unveiled YOLO11 (o…

Computational Efficiencyobject-detectionObject DetectionReal-Time Object Detection