paper-with-me

홈 › Papers

REXUP: I REason, I EXtract, I UPdate with Structured Compositional Reasoning for Visual Question Answering

2020-07-27 · Siwen Luo, Soyeon Caren Han, Kaiyuan Sun, Josiah Poon

Visual question answering (VQA) is a challenging multi-modal task that requires not only the semantic understanding of both images and questions, but also the sound perception of a step-by-step reasoning process that would lead to the correct answer. So far, most successful attempts in VQA have been focused on only one aspect, either the interaction of visual pixel features of images and word features of questions, or the reasoning process of answering the question in an image with simple objects. In this paper, we propose a deep reasoning VQA model with explicit visual structure-aware textual information, and it works well in capturing step-by-step reasoning process and detecting a complex object-relationship in photo-realistic images. REXUP network consists of two branches, image object-oriented and scene graph oriented, which jointly works with super-diagonal fusion compositional attention network. We quantitatively and qualitatively evaluate REXUP on the GQA dataset and conduct extensive ablation studies to explore the reasons behind REXUP's effectiveness. Our best model significantly outperforms the precious state-of-the-art, which delivers 92.7% on the validation set and 73.1% on the test-dev set.

📄 PDF Abstract BibTeX arXiv:2007.13262

Code (1)

usydnlp/REXUP 공식 구현 tf

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Leveraging Textual Compositional Reasoning for Robust Change Captioning

2025-11-28 · Kyu Ri Park, Jiyoung Park, Seong Tae Kim, Hong Joo Lee 외 arxiv

Change captioning aims to describe changes between a pair of images. However, existing works rely on visual features alone, which often fail to capture subtle but meaningful changes because they lack the ability to repre…

Relational Reasoning

Compositional Belief Update

2014-01-15 · James Delgrande, Yi Jin, Francis Jeffry Pelletier

In this paper we explore a class of belief update operators, in which the definition of the operator is compositional with respect to the sentence to be added. The goal is to provide an update operator that is intuitive,…

Sentence

Dynamic MOdularized Reasoning for Compositional Structured Explanation Generation

2023-09-14 · Xiyan Fu, Anette Frank

Despite the success of neural models in solving reasoning tasks, their compositional generalization capabilities remain unclear. In this work, we propose a new setting of the structured explanation generation task to fac…

Explanation Generation

GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning

2023-11-09 · Guangyue Xu, Joyce Chai, Parisa Kordjamshidi

Pre-trained vision-language models (VLMs) have achieved promising success in many fields, especially with prompt learning paradigm. In this work, we propose GIP-COL (Graph-Injected Soft Prompting for COmpositional Learni…

AttributeCompositional Zero-Shot LearningPrompt LearningZero-Shot Learning

CSyMR: Benchmarking Compositional Music Information Retrieval in Symbolic Music Reasoning

2025-12-16 · Boyang Wang, Yash Vishe, Xin Xu, Zachary Novack 외 arxiv

Natural language information needs over symbolic music scores rarely reduce to a single step lookup. Many queries require compositional Music Information Retrieval (MIR) that extracts multiple pieces of evidence from str…

Information Retrieval