paper-with-me

홈 › Papers

Image Captioning with Object Detection and Localization

2017-06-08 · Zhongliang Yang, Yu-Jin Zhang, Sadaqat ur Rehman, Yongfeng Huang

Automatically generating a natural language description of an image is a task close to the heart of image understanding. In this paper, we present a multi-model neural network method closely related to the human visual system that automatically learns to describe the content of images. Our model consists of two sub-models: an object detection and localization model, which extract the information of objects and their spatial relationship in images respectively; Besides, a deep recurrent neural network (RNN) based on long short-term memory (LSTM) units with attention mechanism for sentences generation. Each word of the description will be automatically aligned to different objects of the input image when it is generated. This is similar to the attention mechanism of the human visual system. Experimental results on the COCO dataset showcase the merit of the proposed method, which outperforms previous benchmark models.

📄 PDF Abstract BibTeX arXiv:1706.02430

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningObjectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

DenseCap: Fully Convolutional Localization Networks for Dense Captioning

2015-11-24 · CVPR 2016 6 · Justin Johnson, Andrej Karpathy, Li Fei-Fei

We introduce the dense captioning task, which requires a computer vision system to both localize and describe salient regions in images in natural language. The dense captioning task generalizes object detection when the…

Dense CaptioningImage CaptioningLanguage ModelingLanguage Modelling+3

Transformer based Multitask Learning for Image Captioning and Object Detection

2024-03-10 · Debolena Basak, P. K. Srijith, Maunendra Sankar Desarkar

In several real-world scenarios like autonomous navigation and mobility, to obtain a better visual understanding of the surroundings, image captioning and object detection play a crucial role. This work introduces a nove…

Autonomous NavigationImage CaptioningObjectobject-detection+1

GLIPv2: Unifying Localization and Vision-Language Understanding

2022-06-12 · Haotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen 외

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2…

2D Object DetectionContrastive LearningImage CaptioningInstance Segmentation+10

Spatio-Temporal Attention Models for Grounded Video Captioning

2016-10-17 · Mihai Zanfir, Elisabeta Marinoiu, Cristian Sminchisescu

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and ge…

image-classificationImage ClassificationTemporal LocalizationVideo Captioning

Pixel-Aligned Language Model

2024-01-01 · CVPR 2024 1 · Jiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu 외

Large language models have achieved great success in recent years so as their variants in vision. Existing vision-language models can describe images in natural languages answer visual-related questions or perform co…

Language ModelingLanguage Modellingmodel