paper-with-me

Papers

Multilingual Augmentation for Robust Visual Question Answering in Remote Sensing Images

2023-04-07 · Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu

Aiming at answering questions based on the content of remotely sensed images, visual question answering for remote sensing data (RSVQA) has attracted much attention nowadays. However, previous works in RSVQA have focused little on the robustness of RSVQA. As we aim to enhance the reliability of RSVQA models, how to learn robust representations against new words and different question templates with the same meaning is the key challenge. With the proposed augmented dataset, we are able to obtain more questions in addition to the original ones with the same meaning. To make better use of this information, in this study, we propose a contrastive learning strategy for training robust RSVQA models against diverse question templates and words. Experimental results demonstrate that the proposed augmented dataset is effective in improving the robustness of the RSVQA model. In addition, the contrastive learning strategy performs well on the low resolution (LR) dataset.

📄 PDF Abstract BibTeX arXiv:2304.03844

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningQuestion AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

EVJVQA Challenge: Multilingual Visual Question Answering

2023-02-23 · Ngan Luu-Thuy Nguyen, Nghia Hieu Nguyen, Duong T. D Vo, Khanh Quoc Tran 외

Visual Question Answering (VQA) is a challenging task of natural language processing (NLP) and computer vision (CV), attracting significant attention from researchers. English is a resource-rich language that has witness…

Language ModelingLanguage ModellingQuestion AnsweringVietnamese Multimodal Learning+3

Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images

2024-07-11 · Lucrezia Tosato, Hichem Boussaid, Flora Weissgerber, Camille Kurtz 외

Visual Question Answering for Remote Sensing (RSVQA) is a task that aims at answering natural language questions about the content of a remote sensing image. The visual features extraction is therefore an essential step …

Question AnsweringSegmentationVisual Question AnsweringVisual Question Answering (VQA)

MaXM: Towards Multilingual Visual Question Answering

2022-09-12 · Soravit Changpinyo, Linting Xue, Michal Yarom, Ashish V. Thapliyal 외

Visual Question Answering (VQA) has been primarily studied through the lens of the English language. Yet, tackling VQA in other languages in the same manner would require a considerable amount of resources. In this paper…

Question AnsweringTranslationVisual Question AnsweringVisual Question Answering (VQA)

A Unified Framework for Multilingual and Code-Mixed Visual Question Answering

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we propose an effective deep learning framework for multilingual and code- mixed visual question answering. The pro- posed model is capable of predicting answers from the questions in Hindi, English or Cod…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

xGQA: Cross-Lingual Visual Question Answering

2021-09-13 · Findings (ACL) 2022 5 · Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin O. Steitz 외

Recent advances in multimodal vision and language modeling have predominantly focused on the English language, mostly due to the lack of multilingual multimodal datasets to steer modeling efforts. In this work, we addres…

Cross-Lingual TransferLanguage ModelingLanguage ModellingQuestion Answering+3