paper-with-me

홈 › Papers

Structure-Encoding Auxiliary Tasks for Improved Visual Representation in Vision-and-Language Navigation

2022-11-20 · Chia-Wen Kuo, Chih-Yao Ma, Judy Hoffman, Zsolt Kira

In Vision-and-Language Navigation (VLN), researchers typically take an image encoder pre-trained on ImageNet without fine-tuning on the environments that the agent will be trained or tested on. However, the distribution shift between the training images from ImageNet and the views in the navigation environments may render the ImageNet pre-trained image encoder suboptimal. Therefore, in this paper, we design a set of structure-encoding auxiliary tasks (SEA) that leverage the data in the navigation environments to pre-train and improve the image encoder. Specifically, we design and customize (1) 3D jigsaw, (2) traversability prediction, and (3) instance classification to pre-train the image encoder. Through rigorous ablations, our SEA pre-trained features are shown to better encode structural information of the scenes, which ImageNet pre-trained features fail to properly encode but is crucial for the target navigation task. The SEA pre-trained features can be easily plugged into existing VLN agents without any tuning. For example, on Test-Unseen environments, the VLN agents combined with our SEA pre-trained features achieve absolute success rate improvement of 12% for Speaker-Follower, 5% for Env-Dropout, and 4% for AuxRN.

📄 PDF Abstract BibTeX arXiv:2211.11116

Code (0)

등록된 구현이 없습니다.

Tasks

Test unseenVision and Language Navigation

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens

2025-08-12 · Alexandre Brown, Glen Berseth arxiv

Visual reinforcement learning policies trained on pixel observations often struggle to generalize when visual conditions change at test time. Object-centric representations are a promising alternative, but most approache…

Reinforcement LearningImage Reconstruction

Enhancing Robot Learning through Learned Human-Attention Feature Maps

2023-08-29 · Daniel Scheuchenstuhl, Stefan Ulmer, Felix Resch, Luigi Berducci 외

Robust and efficient learning remains a challenging problem in robotics, in particular with complex visual inputs. Inspired by human attention mechanism, with which we quickly process complex visual scenes and react to c…

Imitation Learningobject-detectionObject DetectionRepresentation Learning

TEACH: Text Encoding as Curriculum Hints for Scene Text Recognition

2025-08-02 · Xiahan Yang, Hui Zheng arxiv

Scene Text Recognition (STR) remains a challenging task due to complex visual appearances and limited semantic priors. We propose TEACH, a novel training paradigm that injects ground-truth text into the model as auxiliar…

Scene Text Recognition

Vision-based Navigation Using Deep Reinforcement Learning

2019-08-08 · Jonáš Kulhánek, Erik Derner, Tim de Bruin, Robert Babuška

Deep reinforcement learning (RL) has been successfully applied to a variety of game-like environments. However, the application of deep RL to visual navigation with realistic environments is a challenging task. We propos…

Deep Reinforcement LearningEfficient Neural Networkreinforcement-learningReinforcement Learning+2

Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions

2025-08-05 · Jingxuan Wei, Caijun Jia, Qi Chen, Honghao He 외 arxiv

Mathematical geometric reasoning is essential for scientific discovery and educational development, requiring precise logic and rigorous formal verification. While recent advances in Multimodal Large Language Models (MLL…

Multimodal Reasoning