Image Captioning 벤치마크
Image Captioning on Flickr30k Captions test
BLEU-4
- 2014-12-07 — BRNN: BLEU-4 15.7
- 2017-06-26 — Cornia et al: BLEU-4 21.3
- 2019-09-24 — Unified VLP: BLEU-4 30.1
| Rank | Model | BLEU-4 | CIDEr | METEOR | SPICE | Extra Training Data | Paper | Code | Year |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Unified VLP | 30.1 | 67.4 | 23 | 17 | Unified Vision-Language Pre-Training for Image Captioning and VQA | rmokady/clip_prefix_caption · LuoweiZhou/VLP · WebQnA/WebQA_Baseline | 2019 | |
| 2 | Cornia et al | 21.3 | 46.4 | 20.0 | - | ✓ | Paying More Attention to Saliency: Image Captioning with Saliency and Context Attention | 2017 | |
| 3 | BRNN | 15.7 | 24.7 | 15.3 | - | Deep Visual-Semantic Alignments for Generating Image Descriptions | VinitSR7/Image-Caption-Generation · Lieberk/Paddle-AoA-Captioning · souvikshanku/digit-captioning · +1 | 2014 | |
| 4 | KOSMOS-1 1.6B (zero-shot) | – | 67.1 | – | 14.5 | ||||
| 5 | MetaLM | – | 43.3 | – | 11.7 | Language Models are General-Purpose Interfaces | microsoft/unilm | 2022 | |
| 6 | FewVLM | – | 31.0 | – | 10.0 | A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models | woojeongjin/fewvlm | 2021 | |
| 7 | VL-T5 | – | 2.6 | – | 2.0 | Unifying Vision-and-Language Tasks via Text Generation | j-min/VL-T5 · mitvis/vistext | 2021 |