paper-with-me

TextCaps

홈페이지 · 논문 98편

Contains 145k captions for 28k images. The dataset challenges a model to recognize text, relate it to its visual context, and decide what part of the text to copy or paraphrase, requiring spatial, semantic, and visual reasoning between multiple text tokens and visual entities, such as objects. Source: [TextCaps: a Dataset for Image Captioning with Reading Comprehension](/paper/textcaps-a-dataset-for-image-captioning-with)

ImagesTexts

벤치마크

Image Captioning on TextCaps 2020 결과 11개