Video Captioning 벤치마크
Video Captioning on TVC
| Rank | Model | BLEU-4 | CIDEr | Extra Training Data | Paper | Code | Year |
|---|---|---|---|---|---|---|---|
| 1 | VAST | 19.9 | 74.1 | ✓ | VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset | TXH-mercury/VALOR · txh-mercury/vast | 2023 |
| 2 | COSA | 18.8 | 70.7 | ✓ | COSA: Concatenated Sample Pretrained Vision-Language Foundation Model | txh-mercury/cosa | 2023 |