paper-with-me

홈 › Papers

VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images

2024-05-06 · Anna Penzkofer, Lei Shi, Andreas Bulling

While Vector Symbolic Architectures (VSAs) are promising for modelling spatial cognition, their application is currently limited to artificially generated images and simple spatial queries. We propose VSA4VQA - a novel 4D implementation of VSAs that implements a mental representation of natural images for the challenging task of Visual Question Answering (VQA). VSA4VQA is the first model to scale a VSA to complex spatial queries. Our method is based on the Semantic Pointer Architecture (SPA) to encode objects in a hyperdimensional vector space. To encode natural images, we extend the SPA to include dimensions for object's width and height in addition to their spatial location. To perform spatial queries we further introduce learned spatial query masks and integrate a pre-trained vision-language model for answering attribute-related questions. We evaluate our method on the GQA benchmark dataset and show that it can effectively encode natural images, achieving competitive performance to state-of-the-art deep learning methods for zero-shot VQA.

📄 PDF Abstract BibTeX arXiv:2405.03852

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeLanguage ModelingLanguage ModellingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Probabilistic Abduction for Visual Abstract Reasoning via Learning Rules in Vector-symbolic Architectures

2024-01-29 · Michael Hersche, Francesco Di Stefano, Thomas Hofmann, Abu Sebastian 외

Abstract reasoning is a cornerstone of human intelligence, and replicating it with artificial intelligence (AI) presents an ongoing challenge. This study focuses on efficiently solving Raven's progressive matrices (RPM),…

Attribute

Capacity Analysis of Vector Symbolic Architectures

2023-01-24 · Kenneth L. Clarkson, Shashanka Ubaru, Elizabeth Yang

Hyperdimensional computing (HDC) is a biologically-inspired framework which represents symbols with high-dimensional vectors, and uses vector operations to manipulate them. The ensemble of a particular vector space and a…

Dimensionality Reduction

Neuromorphic Visual Scene Understanding with Resonator Networks

2022-08-26 · Alpha Renner, Lazar Supic, Andreea Danielescu, Giacomo Indiveri 외

Analyzing a visual scene by inferring the configuration of a generative model is widely considered the most flexible and generalizable approach to scene understanding. Yet, one major problem is the computational challeng…

Scene UnderstandingTranslation

Systematic Abductive Reasoning via Diverse Relation Representations in Vector-symbolic Architecture

2025-01-21 · Zhong-Hua Sun, Ru-Yuan Zhang, Zonglei Zhen, Da-Hui Wang 외

In abstract visual reasoning, monolithic deep learning models suffer from limited interpretability and generalization, while existing neuro-symbolic approaches fall short in capturing the diversity and systematicity of a…

AttributeDiversityOut-of-Distribution GeneralizationRelation+1

Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning

2024-12-07 · Michael Hersche, Giacomo Camposampiero, Roger Wattenhofer, Abu Sebastian 외

This work compares large language models (LLMs) and neuro-symbolic approaches in solving Raven's progressive matrices (RPM), a visual abstract reasoning test that involves the understanding of mathematical rules such as …

Attribute