paper-with-me

홈 › Papers

Simple Image Description Generator via a Linear Phrase-Based Approach

2014-12-29 · Remi Lebret, Pedro O. Pinheiro, Ronan Collobert

Generating a novel textual description of an image is an interesting problem that connects computer vision and natural language processing. In this paper, we present a simple model that is able to generate descriptive sentences given a sample image. This model has a strong focus on the syntax of the descriptions. We train a purely bilinear model that learns a metric between an image representation (generated from a previously trained Convolutional Neural Network) and phrases that are used to described them. The system is then able to infer phrases from a given image sample. Based on caption syntax statistics, we propose a simple language model that can produce relevant descriptions for a given test image using the phrases inferred. Our approach, which is considerably simpler than state-of-the-art models, achieves comparable results on the recently release Microsoft COCO dataset.

📄 PDF Abstract BibTeX arXiv:1412.8419

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveImage DescriptionLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Phrase-based Image Captioning

2015-02-12 · Rémi Lebret, Pedro O. Pinheiro, Ronan Collobert

Generating a novel textual description of an image is an interesting problem that connects computer vision and natural language processing. In this paper, we present a simple model that is able to generate descriptive se…

DescriptiveImage CaptioningLanguage ModelingLanguage Modelling

Generalised Medical Phrase Grounding

2025-11-30 · Wenjun Zhang, Shekhar S. Chandra, Aaron Nicolson arxiv

Medical phrase grounding (MPG) maps textual descriptions of radiological findings to corresponding image regions. These grounded reports are easier to interpret, especially for non-experts. Existing MPG systems mostly fo…

Referring ExpressionPhrase Grounding

Text-to-Image Generation Grounded by Fine-Grained User Attention

2020-11-07 · Jing Yu Koh, Jason Baldridge, Honglak Lee, Yinfei Yang

Localized Narratives is a dataset with detailed natural language descriptions of images paired with mouse traces that provide a sparse, fine-grained visual grounding for phrases. We propose TReCS, a sequential model that…

Image GenerationPositionRetrievalSegmentation+3

Similar Scenes arouse Similar Emotions: Parallel Data Augmentation for Stylized Image Captioning

2021-08-26 · Guodun Li, Yuchen Zhai, Zehao Lin, Yin Zhang

Stylized image captioning systems aim to generate a caption not only semantically related to a given image but also consistent with a given style description. One of the biggest challenges with this task is the lack of s…

Data AugmentationImage CaptioningSentence

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

2026-07-30 · Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang 외 arxiv

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captionin…

Video ReconstructionVideo Captioning