paper-with-me

홈 › Papers

Overcoming Language Bias in Remote Sensing Visual Question Answering via Adversarial Training

2023-06-01 · Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu

The Visual Question Answering (VQA) system offers a user-friendly interface and enables human-computer interaction. However, VQA models commonly face the challenge of language bias, resulting from the learned superficial correlation between questions and answers. To address this issue, in this study, we present a novel framework to reduce the language bias of the VQA for remote sensing data (RSVQA). Specifically, we add an adversarial branch to the original VQA framework. Based on the adversarial branch, we introduce two regularizers to constrain the training process against language bias. Furthermore, to evaluate the performance in terms of language bias, we propose a new metric that combines standard accuracy with the performance drop when incorporating question and random image information. Experimental results demonstrate the effectiveness of our method. We believe that our method can shed light on future work for reducing language bias on the RSVQA task.

📄 PDF Abstract BibTeX arXiv:2306.00483

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation

2023-11-28 · Christel Chappuis, Eliot Walt, Vincent Mendez, Sylvain Lobry 외

Remote sensing visual question answering (RSVQA) opens new opportunities for the use of overhead imagery by the general public, by enabling human-machine interaction with natural language. Building on the recent advances…

DiversityQuestion AnsweringQuestion GenerationQuestion-Generation+2

Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models

2024-01-17 · HaoNan Guo, Xin Su, Chen Wu, Bo Du 외

Recently, the flourishing large language models(LLM), especially ChatGPT, have shown exceptional performance in language understanding, reasoning, and interaction, attracting users and researchers from multiple fields an…

Task Planning

FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing

2025-12-30 · Yunkai Dang, Donghao Wang, Jiacheng Yang, Yifan Jiang 외 arxiv

Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain due to the inherent differences between …

Image Captioning

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

2026-05-18 · Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan 외 arxiv

Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoning. While effective for scene-level understanding, this pipeline may…

Spatial Reasoning

Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning

2023-12-02 · Cong Yang, Zuchao Li, Lefei Zhang

Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this fiel…

Causal Language ModelingContrastive LearningImage CaptioningLanguage Modeling+3