paper-with-me

Papers

Learning to Stop: A Simple yet Effective Approach to Urban Vision-Language Navigation

2020-09-28 · Findings of the Association for Computational Linguistics 2020 · Jiannan Xiang, Xin Eric Wang, William Yang Wang

Vision-and-Language Navigation (VLN) is a natural language grounding task where an agent learns to follow language instructions and navigate to specified destinations in real-world environments. A key challenge is to recognize and stop at the correct location, especially for complicated outdoor environments. Existing methods treat the STOP action equally as other actions, which results in undesirable behaviors that the agent often fails to stop at the destination even though it might be on the right path. Therefore, we propose Learning to Stop (L2Stop), a simple yet effective policy module that differentiates STOP and other actions. Our approach achieves the new state of the art on a challenging urban VLN dataset Touchdown, outperforming the baseline by 6.89% (absolute improvement) on Success weighted by Edit Distance (SED).

📄 PDF Abstract BibTeX arXiv:2009.13112

Code (0)

등록된 구현이 없습니다.

Tasks

NavigateVision and Language NavigationVision-Language Navigation

Similar Papers 제목 키워드 기반

Towards a text-based quantitative and explainable histopathology image analysis

2024-07-10 · Anh Tien Nguyen, Trinh Thi Le Vuong, Jin Tae Kwak

Recently, vision-language pre-trained models have emerged in computational pathology. Previous works generally focused on the alignment of image-text pairs via the contrastive pre-training paradigm. Such pre-trained mode…

image-classificationImage ClassificationImage to textImage-to-Text Retrieval+5

DualProtoSeg: Simple and Efficient Design with Text- and Image-Guided Prototype Learning for Weakly Supervised Histopathology Image Segmentation

2025-12-11 · Anh M. Vu, Khang P. Le, Trang T. K. Vo, Ha Thach 외 arxiv

Weakly supervised semantic segmentation (WSSS) in histopathology seeks to reduce annotation cost by learning from image-level labels, yet it remains limited by inter-class homogeneity, intra-class heterogeneity, and the …

Semantic SegmentationImage Segmentation

Beautimeter: Harnessing GPT for Assessing Architectural and Urban Beauty based on the 15 Properties of Living Structure

2024-11-28 · Bin Jiang

Beautimeter is a new tool powered by generative pre-trained transformer (GPT) technology, designed to evaluate architectural and urban beauty. Rooted in Christopher Alexander's theory of centers, this work builds on the …

Towards Zero-Shot Annotation of the Built Environment with Vision-Language Models (Vision Paper)

2024-08-01 · Bin Han, Yiwei Yang, Anat Caspi, Bill Howe

Equitable urban transportation applications require high-fidelity digital representations of the built environment: not just streets and sidewalks, but bike lanes, marked and unmarked crossings, curb ramps and cuts, obst…

Language Modelling

Unsupervised Abnormal Stop Detection for Long Distance Coaches with Low-Frequency GPS

2024-11-07 · Jiaxin Deng, Junbiao Pang, Jiayu Xu, HaiTao Yu

In our urban life, long distance coaches supply a convenient yet economic approach to the transportation of the public. One notable problem is to discover the abnormal stop of the coaches due to the important reason, i.e…