paper-with-me

홈 › Papers

Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach

2024-05-01 · Zhilin Zhang, Fangyu Wu

Visual Question Answering (VQA) has emerged as a highly engaging field in recent years, with increasing research focused on enhancing VQA accuracy through advanced models such as Transformers. Despite this growing interest, limited work has examined the comparative effectiveness of textual encoders in VQA, particularly considering model complexity and computational efficiency. In this work, we conduct a comprehensive comparison between complex textual models that leverage long-range dependencies and simpler models focusing on local textual features within a well-established VQA framework. Our findings reveal that employing complex textual encoders is not always the optimal approach for the VQA-v2 dataset. Motivated by this insight, we propose ConvGRU, a model that incorporates convolutional layers to improve text feature representation without substantially increasing model complexity. Tested on the VQA-v2 dataset, ConvGRU demonstrates a modest yet consistent improvement over baselines for question types such as Number and Count, which highlights the potential of lightweight architectures for VQA tasks, especially when computational resources are limited.

📄 PDF Abstract BibTeX arXiv:2405.00479

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

A Dual-Attention Learning Network with Word and Sentence Embedding for Medical Visual Question Answering

2022-10-01 · Xiaofei Huang, Hongfang Gong

Research in medical visual question answering (MVQA) can contribute to the development of computeraided diagnosis. MVQA is a task that aims to predict accurate and convincing answers based on given medical images and ass…

Medical Visual Question AnsweringQuestion AnsweringSentenceSentence Embedding+4

Good Visual Guidance Make A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction

2022-07-01 · Findings (NAACL) 2022 7 · Xiang Chen, Ningyu Zhang, Lei LI, Yunzhi Yao 외

Multimodal named entity recognition and relation extraction (MNER and MRE) is a fundamental and crucial branch in information extraction. However, existing approaches for MNER and MRE usually suffer from error sensitivit…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Relation+1

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

2019-04-08 · CVPR 2019 6 · Chenyou Fan, Xiaofan Zhang, Shu Zhang, Wensheng Wang 외

In this paper, we propose a novel end-to-end trainable Video Question Answering (VideoQA) framework with three major components: 1) a new heterogeneous memory which can effectively learn global context information from a…

Question AnsweringVideo Question AnsweringVisual Question Answering (VQA)

BLCU\_NLP at SemEval-2019 Task 8: A Contextual Knowledge-enhanced GPT Model for Fact Checking

2019-06-01 · SEMEVAL 2019 6 · Wanying Xie, Mengxi Que, Ruoyao Yang, Chunhua Liu 외

Since the resources of Community Question Answering are abundant and information sharing becomes universal, it will be increasingly difficult to find factual information for questioners in massive messages. SemEval 2019 …

Community Question AnsweringFact CheckingQuestion Answering

Good Visual Guidance Makes A Better Extractor: Hierarchical Visual Prefix for Multimodal Entity and Relation Extraction

2022-05-07 · Xiang Chen, Ningyu Zhang, Lei LI, Yunzhi Yao 외

Multimodal named entity recognition and relation extraction (MNER and MRE) is a fundamental and crucial branch in information extraction. However, existing approaches for MNER and MRE usually suffer from error sensitivit…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Relation+1