From Captions to Visual Concepts and Back
This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.
Code (1)
Tasks
Image CaptioningLanguage ModelingLanguage ModellingMultiple Instance LearningRe-RankingSentenceSimilar Papers 제목 키워드 기반
TCIC: Theme Concepts Learning Cross Language and Vision for Image Captioning
Existing research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concept…
DecoderImage CaptioningRepresentation LearningCultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
Pretrained vision-language models (VLMs) such as CLIP excel in multimodal understanding but struggle with contextually relevant fine-grained visual features, making it difficult to distinguish visually similar yet cultur…
Contrastive LearningHyperbolic Learning with Synthetic Captions for Open-World Detection
Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption dat…
HallucinationNovel ConceptsObjectobject-detection+1A Semi-supervised Framework for Image Captioning
State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data. These methods are typically unable to use unsupervised data such as textual data with no co…
DecoderImage CaptioningWord Embeddingsnocaps: novel object captioning at scale
Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, …
Image CaptioningObjectobject-detectionObject Detection