Vision and Language Navigation 벤치마크
Vision and Language Navigation on RxR
ndtw
- 2020-10-15 — Monolingual Baseline: ndtw 41.05
- 2021-07-13 — CLEAR-CLIP: ndtw 53.69
- 2021-10-25 — HAMT: ndtw 59.94
- 2022-03-29 — EnvEdit-PT: ndtw 64.61
- 2022-10-06 — MARVAL: ndtw 66.76
| Rank | Model | ndtw | Extra Training Data | Paper | Code | Year |
|---|---|---|---|---|---|---|
| 1 | MARVAL | 66.76 | ✓ | A New Path: Scaling Vision-and-Language Navigation with Synthetic Instructions and Imitation Learning | 2022 | |
| 2 | EnvEdit-PT | 64.61 | ✓ | EnvEdit: Environment Editing for Vision-and-Language Navigation | jialuli-luka/envedit | 2022 |
| 3 | HAMT | 59.94 | History Aware Multimodal Transformer for Vision-and-Language Navigation | cshizhe/vln-hamt | 2021 | |
| 4 | CLEAR-CLIP | 53.69 | ✓ | How Much Can CLIP Benefit Vision-and-Language Tasks? | clip-vil/CLIP-ViL · jianjieluo/openai-clip-feature · facebookresearch/reliable_vqa · +1 | 2021 |
| 5 | Monolingual Baseline | 41.05 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding | jacobkrantz/VLN-CE · google-research-datasets/RxR · VegB/Diagnose_VLN | 2020 | |
| 6 | Multilingual Baseline | 36.81 | Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding | jacobkrantz/VLN-CE · google-research-datasets/RxR · VegB/Diagnose_VLN | 2020 |