paper-with-me

Papers

Image Captioning and Visual Question Answering Based on Attributes and External Knowledge

2016-03-09 · Qi Wu, Chunhua Shen, Anton Van Den Hengel, Peng Wang, Anthony Dick

Much recent progress in Vision-to-Language problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image features to text. In this paper we first propose a method of incorporating high-level concepts into the successful CNN-RNN approach, and show that it achieves a significant improvement on the state-of-the-art in both image captioning and visual question answering. We further show that the same mechanism can be used to incorporate external knowledge, which is critically important for answering high level visual questions. Specifically, we design a visual question answering model that combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. It particularly allows questions to be asked about the contents of an image, even when the image itself does not contain a complete answer. Our final model achieves the best reported results on both image captioning and visual question answering on several benchmark datasets.

📄 PDF Abstract BibTeX arXiv:1603.02814

Code (0)

등록된 구현이 없습니다.

Tasks

General KnowledgeImage CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

ParsVQA-Caps: A Benchmark for Visual Question Answering and Image Captioning in Persian

2022-12-07 · WiNLP2022 2022 12 · Shaghayegh Mobasher, Ghazal Zamaninejad, Maryam Hashemi, Melika Nobakhtian 외

Despite recent advances in vision-and-language tasks, most progress is still focused on resource-rich languages such as English. Furthermore, widespread vision-and-language datasets directly adopt images representative o…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering

2019-09-04 · IJCNLP 2019 11 · Soravit Changpinyo, Bo Pang, Piyush Sharma, Radu Soricut

Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster R-CNN rely on a costly process of annota…

Image CaptioningObjectobject-detectionObject Detection+4

Is GPT-3 all you need for Visual Question Answering in Cultural Heritage?

2022-07-25 · Pietro Bongini, Federico Becattini, Alberto del Bimbo

The use of Deep Learning and Computer Vision in the Cultural Heritage domain is becoming highly relevant in the last few years with lots of applications about audio smart guides, interactive museums and augmented reality…

AllQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Joint Image Captioning and Question Answering

2018-05-22 · Jialin Wu, Zeyuan Hu, Raymond J. Mooney

Answering visual questions need acquire daily common knowledge and model the semantic connection among different parts in images, which is too difficult for VQA systems to learn from images with the only supervision from…

Image CaptioningQuestion AnsweringVisual Question Answering (VQA)

Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts

2024-04-12 · Övgü Özdemir, Erdem Akagündüz

Visual question answering (VQA) is known as an AI-complete task as it requires understanding, reasoning, and inferring about the vision and the language content. Over the past few years, numerous neural architectures hav…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)