Image Captioning with Object Detection and Localization
Automatically generating a natural language description of an image is a task close to the heart of image understanding. In this paper, we present a multi-model neural network method closely related to the human visual system that automatically learns to describe the content of images. Our model consists of two sub-models: an object detection and localization model, which extract the information of objects and their spatial relationship in images respectively; Besides, a deep recurrent neural network (RNN) based on long short-term memory (LSTM) units with attention mechanism for sentences generation. Each word of the description will be automatically aligned to different objects of the input image when it is generated. This is similar to the attention mechanism of the human visual system. Experimental results on the COCO dataset showcase the merit of the proposed method, which outperforms previous benchmark models.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningObjectobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
DenseCap: Fully Convolutional Localization Networks for Dense Captioning
We introduce the dense captioning task, which requires a computer vision system to both localize and describe salient regions in images in natural language. The dense captioning task generalizes object detection when the…
Dense CaptioningImage CaptioningLanguage ModelingLanguage Modelling+3Transformer based Multitask Learning for Image Captioning and Object Detection
In several real-world scenarios like autonomous navigation and mobility, to obtain a better visual understanding of the surroundings, image captioning and object detection play a crucial role. This work introduces a nove…
Autonomous NavigationImage CaptioningObjectobject-detection+1GLIPv2: Unifying Localization and Vision-Language Understanding
We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2…
2D Object DetectionContrastive LearningImage CaptioningInstance Segmentation+10Spatio-Temporal Attention Models for Grounded Video Captioning
Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and ge…
image-classificationImage ClassificationTemporal LocalizationVideo CaptioningPixel-Aligned Language Model
Large language models have achieved great success in recent years so as their variants in vision. Existing vision-language models can describe images in natural languages answer visual-related questions or perform co…
Language ModelingLanguage Modellingmodel