Cross-Modal Retrieval on Flickr30k
Image-to-text R@1
- 2017-07-18 — VSE++ (ResNet): Image-to-text R@1 52.9
- 2017-11-15 — Dual-Path (ResNet): Image-to-text R@1 55.6
- 2018-03-21 — SCAN: Image-to-text R@1 67.4
- 2020-03-08 — IMRAM: Image-to-text R@1 74.1
- 2020-04-01 — GSMN: Image-to-text R@1 76.4
- 2021-01-05 — SGRAF: Image-to-text R@1 77.8
- 2021-02-05 — ViLT-B/32: Image-to-text R@1 83.5
- 2021-02-11 — ALIGN: Image-to-text R@1 95.3
- 2021-11-16 — X-VLM (base): Image-to-text R@1 97.1
- 2022-08-22 — BEiT-3: Image-to-text R@1 98.0
- 2022-11-22 — X2-VLM (large): Image-to-text R@1 98.8