paper-with-me

홈 › Papers

From Captions to Visual Concepts and Back

2014-11-18 · CVPR 2015 6 · Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C. Platt, C. Lawrence Zitnick, Geoffrey Zweig

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.

📄 PDF Abstract BibTeX arXiv:1411.4952

Code (1)

s-gupta/visual-concepts 공식 구현

Tasks

Image CaptioningLanguage ModelingLanguage ModellingMultiple Instance LearningRe-RankingSentence

Similar Papers 제목 키워드 기반

TCIC: Theme Concepts Learning Cross Language and Vision for Image Captioning

2021-06-21 · Zhihao Fan, Zhongyu Wei, Siyuan Wang, Ruize Wang 외

Existing research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concept…

DecoderImage CaptioningRepresentation Learning

CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions

2025-07-08 · Yuchen Huang, Zhiyuan Fan, Zhitao He, Sandeep Polisetty 외

Pretrained vision-language models (VLMs) such as CLIP excel in multimodal understanding but struggle with contextually relevant fine-grained visual features, making it difficult to distinguish visually similar yet cultur…

Contrastive Learning

Hyperbolic Learning with Synthetic Captions for Open-World Detection

2024-04-07 · CVPR 2024 1 · Fanjie Kong, Yanbei Chen, Jiarui Cai, Davide Modolo

Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption dat…

HallucinationNovel ConceptsObjectobject-detection+1

A Semi-supervised Framework for Image Captioning

2016-11-16 · Wenhu Chen, Aurelien Lucchi, Thomas Hofmann

State-of-the-art approaches for image captioning require supervised training data consisting of captions with paired image data. These methods are typically unable to use unsupervised data such as textual data with no co…

DecoderImage CaptioningWord Embeddings

nocaps: novel object captioning at scale

2018-12-20 · ICCV 2019 10 · Harsh Agrawal, Karan Desai, YuFei Wang, Xinlei Chen 외

Image captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, …

Image CaptioningObjectobject-detectionObject Detection