paper-with-me

홈 › Papers

Question-Led Semantic Structure Enhanced Attentions for VQA

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The exploit of the semantic structure in the visual question answering (VQA) task is a trending topic where researchers are interested in leveraging internal semantics and bringing in external knowledge to tackle more complex questions. The prevailing approaches either encode the external knowledge separately from the local context, which magnificently increases the complexity of the ensemble system, or use graph neural networks to model the semantic structure in the context, which suffers from the limited reasoning capability due to the relatively shallow network. In this work, we propose a question-led structure extraction scheme using external knowledge and explore multiple training methods, including direct attention supervision, SGHMC-EM Bayesian multitask learning, and masking strategies, to aggregate the structural knowledge into deep models without changing the architectures. We conduct extensive experiments on two domain-specific but challenging sub-tasks of VrR-VG dataset and demonstrate that our proposed methods achieve significant improvements over strong baselines, showing the promising potentials of applicability.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Inverse Visual Question Answering with Multi-Level Attentions

2019-09-17 · Yaser Alwattar, Yuhong Guo

In this paper, we propose a novel deep multi-level attention model to address inverse visual question answering. The proposed model generates regional visual and semantic features at the object level and then enhances th…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

A Better Way to Attend: Attention with Trees for Video Question Answering

2019-09-05 · Hongyang Xue, Wenqing Chu, Zhou Zhao, Deng Cai

We propose a new attention model for video question answering. The main idea of the attention models is to locate on the most informative parts of the visual data. The attention mechanisms are quite popular these days. H…

Question AnsweringVideo Question Answering

IDEA: Interactive DoublE Attentions from Label Embedding for Text Classification

2022-09-23 · Ziyuan Wang, Hailiang Huang, Songqiao Han

Current text classification methods typically encode the text merely into embedding before a naive or complicated classifier, which ignores the suggestive information contained in the label text. As a matter of fact, hum…

text-classificationText Classification

Dense but Efficient VideoQA for Intricate Compositional Reasoning

2022-10-19 · Jihyeon Lee, Wooyoung Kang, Eun-Sol Kim

It is well known that most of the conventional video question answering (VideoQA) datasets consist of easy questions requiring simple reasoning processes. However, long videos inevitably contain complex and compositional…

Question AnsweringVideo Question Answering

All-In-One Metrical And Functional Structure Analysis With Neighborhood Attentions on Demixed Audio

2023-07-31 · Taejun Kim, Juhan Nam

Music is characterized by complex hierarchical structures. Developing a comprehensive model to capture these structures has been a significant challenge in the field of Music Information Retrieval (MIR). Prior research h…

AllDownbeat TrackingInformation RetrievalMusic Information Retrieval+1