paper-with-me

Papers

Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation

2018-11-25 · CVPR 2019 6 · Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang

Vision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%).

📄 PDF Abstract BibTeX arXiv:1811.10092

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningReinforcement LearningReinforcement Learning (RL)Vision-Language NavigationVisual Navigation

Similar Papers 제목 키워드 기반

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

2026-05-27 · Runlong Cao, Ying Zang, Chuanwei Zhou, Tianrun Chen 외 arxiv

Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers from limited supervision and unreliable pseudo-labels when exploitin…

Referring Expression Segmentation

Self-reinforcing Unsupervised Matching

2019-08-23 · Jiang Lu, Lei LI, Chang-Shui Zhang

Remarkable gains in deep learning usually rely on tremendous supervised data. Ensuring the modality diversity for one object in training set is critical for the generalization of cutting-edge deep models, but it burdens …

Continual LearningDiversityspeech-recognitionSpeech Recognition

Self-Supervised Learning for Multimodal Non-Rigid 3D Shape Matching

2023-03-20 · CVPR 2023 1 · Dongliang Cao, Florian Bernard

The matching of 3D shapes has been extensively studied for shapes represented as surface meshes, as well as for shapes represented as point clouds. While point clouds are a common representation of raw real-world 3D data…

Self-Supervised Learning

Reinforced Cross-modal Alignment for Radiology Report Generation

2022-05-01 · Findings (ACL) 2022 5 · Han Qin, Yan Song

Medical images are widely used in clinical decision-making, where writing radiology reports is a potential application that can be enhanced by automatic solutions to alleviate physicians’ workload. In general, radiology …

cross-modal alignmentDecision MakingReinforcement Learning (RL)valid

Transferring Pre-trained Multimodal Representations with Cross-modal Similarity Matching

2023-01-07 · Byoungjip Kim, Sungik Choi, Dasol Hwang, Moontae Lee 외

Despite surprising performance on zero-shot transfer, pre-training a large-scale multimodal model is often prohibitive as it requires a huge amount of data and computing resources. In this paper, we propose a method (Bea…

Language ModelingLanguage ModellingSelf-Supervised Learning