paper-with-me

Papers

VIBIKNet: Visual Bidirectional Kernelized Network for Visual Question Answering

2016-12-12 · Marc Bolaños, Álvaro Peris, Francisco Casacuberta, Petia Radeva

In this paper, we address the problem of visual question answering by proposing a novel model, called VIBIKNet. Our model is based on integrating Kernelized Convolutional Neural Networks and Long-Short Term Memory units to generate an answer given a question about an image. We prove that VIBIKNet is an optimal trade-off between accuracy and computational load, in terms of memory and time consumption. We validate our method on the VQA challenge dataset and compare it to the top performing methods in order to illustrate its performance and speed.

📄 PDF Abstract BibTeX arXiv:1612.03628

Code (1)

MarcBS/VIBIKNet 공식 구현

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

The Kernelized Taylor Diagram

2022-05-18 · Kristoffer Wickstrøm, J. Emmanuel Johnson, Sigurd Løkse, Gustau Camps-Valls 외

This paper presents the kernelized Taylor diagram, a graphical framework for visualizing similarities between data populations. The kernelized Taylor diagram builds on the widely used Taylor diagram, which is used to vis…

Data Visualization

Write a Classifier: Predicting Visual Classifiers from Unstructured Text

2015-12-31 · Mohamed Elhoseiny, Ahmed Elgammal, Babak Saleh

People typically learn through exposure to visual concepts associated with linguistic descriptions. For instance, teaching visual object categories to children is often accompanied by descriptions in text or speech. In a…

regressionTransfer Learning

Alignment Distances on Systems of Bags

2017-06-14 · Alexander Sagel, Martin Kleinsteuber

Recent research in image and video recognition indicates that many visual processes can be thought of as being generated by a time-varying generative model. A nearby descriptive model for visual processes is thus a stati…

DescriptiveDictionary LearningGeneral ClassificationVideo Recognition

Multi-Level Attention Networks for Visual Question Answering

2017-07-01 · CVPR 2017 7 · Dongfei Yu, Jianlong Fu, Tao Mei, Yong Rui

Inspired by the recent success of text-based question answering, visual question answering (VQA) is proposed to automatically answer natural language questions with the reference to a given image. Compared with text-base…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Task-driven Visual Saliency and Attention-based Visual Question Answering

2017-02-22 · Yuetan Lin, Zhangyang Pang, Donghui Wang, Yueting Zhuang

Visual question answering (VQA) has witnessed great progress since May, 2015 as a classic problem unifying visual and textual data into a system. Many enlightening VQA works explore deep into the image and question encod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)