Captioning Images with Novel Objects via Online Vocabulary Expansion
In this study, we introduce a low cost method for generating descriptions from images containing novel objects. Generally, constructing a model, which can explain images with novel objects, is costly because of the following: (1) collecting a large amount of data for each category, and (2) retraining the entire system. If humans see a small number of novel objects, they are able to estimate their properties by associating their appearance with known objects. Accordingly, we propose a method that can explain images with novel objects without retraining using the word embeddings of the objects estimated from only a small number of image features of the objects. The method can be integrated with general image-captioning models. The experimental results show the effectiveness of our approach.
Code (0)
등록된 구현이 없습니다.
Tasks
Image CaptioningWord EmbeddingsSimilar Papers 제목 키워드 기반
Guided Open Vocabulary Image Captioning with Constrained Beam Search
Existing image captioning models do not generalize well to out-of-domain images containing novel scenes or objects. This limitation severely hinders the use of these models in real world applications dealing with images …
Image CaptioningTAGWord EmbeddingsPointing Novel Objects in Image Captioning
Image captioning has received significant attention with remarkable improvements in recent advances. Nevertheless, images in the wild encapsulate rich knowledge and cannot be sufficiently described with models built on i…
DecoderImage CaptioningObjectObject Recognition+1Auto-Vocabulary 3D Object Detection
Open-vocabulary 3D object detection methods are able to localize 3D boxes of classes unseen during training. Despite the name, existing methods rely on user-specified classes both at training and inference. We propose to…
3D Object DetectionImage CaptioningNOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge
Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection m…
Caption GenerationObjectobject-detectionObject Detection+1Good News, Everyone! Context driven entity-aware captioning for news images
Current image captioning systems perform at a merely descriptive level, essentially enumerating the objects in the scene and their relations. Humans, on the contrary, interpret images by integrating several sources of pr…
ArticlesDescriptiveImage Captioning