paper-with-me

홈 › Papers

Microsoft COCO Captions: Data Collection and Evaluation Server

2015-04-01 · Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollar, C. Lawrence Zitnick

In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.

📄 PDF Abstract BibTeX arXiv:1504.00325

Code (18)

tylin/coco-caption 공식 구현
Chloejay/image_caption_app tf
NUS-VIP/salicon-evaluation
chldydgh4687/2020-1.VideoCaptioning pytorch
daqingliu/coco-caption
gabrielsantosrv/coco-caption pytorch
jmhessel/pycocoevalcap pytorch
luoweizhou/coco-caption
lvapeab/coco-caption
mtanti/coco-caption
peteanderson80/coco-caption
qingzwang/DiversityMetrics tf
ruotianluo/coco-caption pytorch
salaniz/pycocoevalcap
ttengwang/coco-caption
tuetschek/e2e-metrics
vrama91/coco-caption
wenhuchen/data-to-text-evaluation-metric

Tasks

Caption Generation

Methods 이 논문이 사용한 방법론

Microsoft Support 1-855-535-7109: Unlocking Solutions for Your Tech Needs 설명 없음

Similar Papers 제목 키워드 기반

Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task

2019-10-08 · EMNLP (WS) 2019 11 · Alireza Mohammadshahi, Remi Lebret, Karl Aberer

In this paper, we propose a new approach to learn multimodal multilingual embeddings for matching images and their relevant captions in two languages. We combine two existing objective functions to make images and captio…

Cross-Modal RetrievalImage to textImage-to-Text RetrievalMultilingual Word Embeddings+3

ChatPainter: Improving Text to Image Generation using Dialogue

2018-02-22 · Shikhar Sharma, Dendi Suhubdy, Vincent Michalski, Samira Ebrahimi Kahou 외

Synthesizing realistic images from text descriptions on a dataset like Microsoft Common Objects in Context (MS COCO), where each image can contain several objects, is a challenging task. Prior work has used text captions…

Image GenerationText to Image GenerationText-to-Image Generation

Joint Learning of Distributed Representations for Images and Texts

2015-04-13 · Xiaodong He, Rupesh Srivastava, Jianfeng Gao, Li Deng

This technical report provides extra details of the deep multimodal similarity model (DMSM) which was proposed in (Fang et al. 2015, arXiv:1411.4952). The model is trained via maximizing global semantic similarity betwee…

Semantic SimilaritySemantic Textual Similarity

Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion

2023-06-20 · Simone Bianco, Luigi Celona, Marco Donzella, Paolo Napoletano

State-of-The-Art (SoTA) image captioning models often rely on the Microsoft COCO (MS-COCO) dataset for training. This dataset contains annotations provided by human annotators, who typically produce captions averaging ar…

Image CaptioningLanguage ModellingLarge Language Model

Alleviating Noisy Data in Image Captioning with Cooperative Distillation

2020-12-21 · Pierre Dognin, Igor Melnyk, Youssef Mroueh, Inkit Padhi 외

Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their corresponding images. Unfortunately, sca…

Image Captioning