paper-with-me

Papers

Visual Commonsense based Heterogeneous Graph Contrastive Learning

2023-11-11 · Zongzhao Li, Xiangyu Zhu, Xi Zhang, Zhaoxiang Zhang, Zhen Lei

How to select relevant key objects and reason about the complex relationships cross vision and linguistic domain are two key issues in many multi-modality applications such as visual question answering (VQA). In this work, we incorporate the visual commonsense information and propose a heterogeneous graph contrastive learning method to better finish the visual reasoning task. Our method is designed as a plug-and-play way, so that it can be quickly and easily combined with a wide range of representative methods. Specifically, our model contains two key components: the Commonsense-based Contrastive Learning and the Graph Relation Network. Using contrastive learning, we guide the model concentrate more on discriminative objects and relevant visual commonsense attributes. Besides, thanks to the introduction of the Graph Relation Network, the model reasons about the correlations between homogeneous edges and the similarities between heterogeneous edges, which makes information transmission more effective. Extensive experiments on four benchmarks show that our method greatly improves seven representative VQA models, demonstrating its effectiveness and generalizability.

📄 PDF Abstract BibTeX arXiv:2311.06553

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningQuestion AnsweringRelationRelation NetworkVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Heterogeneous Graph Learning for Visual Commonsense Reasoning

2019-10-25 · NeurIPS 2019 12 · Weijiang Yu, Jingwen Zhou, Weihao Yu, Xiaodan Liang 외

Visual commonsense reasoning task aims at leading the research field into solving cognition-level reasoning with the ability of predicting correct answers and meanwhile providing convincing reasoning paths, resulting in …

Graph LearningVisual Commonsense Reasoning

DIVE: Towards Descriptive and Diverse Visual Commonsense Generation

2024-08-15 · Jun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon 외

Towards human-level visual understanding, visual commonsense generation has been introduced to generate commonsense inferences beyond images. However, current research on visual commonsense generation has overlooked an i…

DescriptiveDiversity

Str-GCL: Structural Commonsense Driven Graph Contrastive Learning

2025-07-09 · Dongxiao He, Yongqi Huang, Jitao Zhao, Xiaobao Wang 외 arxiv

Graph Contrastive Learning (GCL) is a widely adopted approach in self-supervised graph representation learning, applying contrastive objectives to produce effective representations. However, current GCL methods primarily…

Graph Representation LearningContrastive Learning

Expressive Scene Graph Generation Using Commonsense Knowledge Infusion for Visual Understanding and Reasoning

2022-05-31 · European Semantic Web Conference (ESWC) 2022 5 · Khan, M. Jaleed; Breslin, John G.; Curry, Edward

Scene graph generation aims to capture the semantic elements in images by modelling objects and their relationships in a structured manner, which are essential for visual understanding and reasoning tasks including image…

Common Sense ReasoningGraph GenerationImage CaptioningImage Generation+9

Learning from Missing Relations: Contrastive Learning with Commonsense Knowledge Graphs for Commonsense Inference

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Commonsense inference poses a unique challenge to reason and generate the physical, social, and causal conditions of a given event. Existing approaches to commonsense inference utilize commonsense transformers, which are…

Contrastive LearningDiversityKnowledge Graphs