Microsoft COCO Captions: Data Collection and Evaluation Server
In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.
Code (18)
Tasks
Caption GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task
In this paper, we propose a new approach to learn multimodal multilingual embeddings for matching images and their relevant captions in two languages. We combine two existing objective functions to make images and captio…
Cross-Modal RetrievalImage to textImage-to-Text RetrievalMultilingual Word Embeddings+3ChatPainter: Improving Text to Image Generation using Dialogue
Synthesizing realistic images from text descriptions on a dataset like Microsoft Common Objects in Context (MS COCO), where each image can contain several objects, is a challenging task. Prior work has used text captions…
Image GenerationText to Image GenerationText-to-Image GenerationJoint Learning of Distributed Representations for Images and Texts
This technical report provides extra details of the deep multimodal similarity model (DMSM) which was proposed in (Fang et al. 2015, arXiv:1411.4952). The model is trained via maximizing global semantic similarity betwee…
Semantic SimilaritySemantic Textual SimilarityImproving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
State-of-The-Art (SoTA) image captioning models often rely on the Microsoft COCO (MS-COCO) dataset for training. This dataset contains annotations provided by human annotators, who typically produce captions averaging ar…
Image CaptioningLanguage ModellingLarge Language ModelAlleviating Noisy Data in Image Captioning with Cooperative Distillation
Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their corresponding images. Unfortunately, sca…
Image Captioning