paper-with-me

홈 › Papers

SelfGraphVQA: A Self-Supervised Graph Neural Network for Scene-based Question Answering

2023-10-03 · Bruno Souza, Marius Aasan, Helio Pedrini, Adín Ramírez Rivera

The intersection of vision and language is of major interest due to the increased focus on seamless integration between recognition and reasoning. Scene graphs (SGs) have emerged as a useful tool for multimodal image analysis, showing impressive performance in tasks such as Visual Question Answering (VQA). In this work, we demonstrate that despite the effectiveness of scene graphs in VQA tasks, current methods that utilize idealized annotated scene graphs struggle to generalize when using predicted scene graphs extracted from images. To address this issue, we introduce the SelfGraphVQA framework. Our approach extracts a scene graph from an input image using a pre-trained scene graph generator and employs semantically-preserving augmentation with self-supervised techniques. This method improves the utilization of graph representations in VQA tasks by circumventing the need for costly and potentially biased annotated data. By creating alternative views of the extracted graphs through image augmentations, we can learn joint embeddings by optimizing the informational content in their representations using an un-normalized contrastive approach. As we work with SGs, we experiment with three distinct maximization strategies: node-wise, graph-wise, and permutation-equivariant regularization. We empirically showcase the effectiveness of the extracted scene graph for VQA and demonstrate that these approaches enhance overall performance by highlighting the significance of visual information. This offers a more practical solution for VQA tasks that rely on SGs for complex reasoning questions.

📄 PDF Abstract BibTeX arXiv:2310.01842

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Neural NetworkQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Do we still need ImageNet pre-training in remote sensing scene classification?

2021-11-05 · Vladimir Risojević, Vladan Stojnić

Due to the scarcity of labeled data, using supervised models pre-trained on ImageNet is a de facto standard in remote sensing scene classification. Recently, the availability of larger high resolution remote sensing (HRR…

ClassificationMulti-Label ClassificationScene ClassificationSelf-Supervised Learning

Self-Supervised Relation Alignment for Scene Graph Generation

2023-02-02 · Bicheng Xu, Renjie Liao, Leonid Sigal

The goal of scene graph generation is to predict a graph from an input image, where nodes correspond to identified and localized objects and edges to their corresponding interaction predicates. Existing methods are train…

Graph GenerationRelationRelation PredictionScene Graph Generation

Self Supervised Clustering of Traffic Scenes using Graph Representations

2022-11-24 · Maximilian Zipfl, Moritz Jarosch, J. Marius Zöllner

Examining graphs for similarity is a well-known challenge, but one that is mandatory for grouping graphs together. We present a data-driven method to cluster traffic scenes that is self-supervised, i.e. without manual la…

ClusteringGraph Embedding

SGRec3D: Self-Supervised 3D Scene Graph Learning via Object-Level Scene Reconstruction

2023-09-27 · Sebastian Koch, Pedro Hermosilla, Narunas Vaskevicius, Mirco Colosi 외

In the field of 3D scene understanding, 3D scene graphs have emerged as a new scene representation that combines geometric and semantic information about objects and their relationships. However, learning semantic 3D sce…

Graph LearningPredictionScene Understanding

Self-supervised Knowledge Triplet Learning for Zero-shot Question Answering

2020-05-01 · EMNLP 2020 11 · Pratyay Banerjee, Chitta Baral

The aim of all Question Answering (QA) systems is to be able to generalize to unseen questions. Current supervised methods are reliant on expensive data annotation. Moreover, such annotations can introduce unintended ann…

Knowledge GraphsQuestion AnsweringTriplet