paper-with-me

홈 › Papers

Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning

2021-09-14 · EMNLP 2021 11 · Da Yin, Liunian Harold Li, Ziniu Hu, Nanyun Peng, Kai-Wei Chang

Commonsense is defined as the knowledge that is shared by everyone. However, certain types of commonsense knowledge are correlated with culture and geographic locations and they are only shared locally. For example, the scenarios of wedding ceremonies vary across regions due to different customs influenced by historical and religious factors. Such regional characteristics, however, are generally omitted in prior work. In this paper, we construct a Geo-Diverse Visual Commonsense Reasoning dataset (GD-VCR) to test vision-and-language models' ability to understand cultural and geo-location-specific commonsense. In particular, we study two state-of-the-art Vision-and-Language models, VisualBERT and ViLBERT trained on VCR, a standard multimodal commonsense benchmark with images primarily from Western regions. We then evaluate how well the trained models can generalize to answering the questions in GD-VCR. We find that the performance of both models for non-Western regions including East Asia, South Asia, and Africa is significantly lower than that for Western region. We analyze the reasons behind the performance disparity and find that the performance gap is larger on QA pairs that: 1) are concerned with culture-related scenarios, e.g., weddings, religious activities, and festivals; 2) require high-level geo-diverse commonsense reasoning rather than low-order perception and recognition. Dataset and code are released at https://github.com/WadeYin9712/GD-VCR.

📄 PDF Abstract BibTeX arXiv:2109.06860

Code (1)

wadeyin9712/gd-vcr 공식 구현 pytorch

Tasks

Cultural Vocal Bursts Intensity PredictionVisual Commonsense Reasoning

Methods 이 논문이 사용한 방법론

VisualBERT VisualBERT aims to reuse self-attention to implicitly align elements of the input text and regions in the input image. Visual embeddings are used to model images where the…
ViLBERT Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and…

Similar Papers 제목 키워드 기반

Commonsense Visual Sensemaking for Autonomous Driving: On Generalised Neurosymbolic Online Abduction Integrating Vision and Semantics

2020-12-28 · Jakob Suchan, Mehul Bhatt, Srikrishna Varadarajan

We demonstrate the need and potential of systematically integrated vision and semantics solutions for visual sensemaking in the backdrop of autonomous driving. A general neurosymbolic method for online visual sensemaking…

Autonomous DrivingQuestion AnsweringSpatial Reasoning

Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs

2025-10-21 · Jiaao Yu, Shenwei Li, Mingjie Han, Yifei Yin 외 arxiv

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their ad…

Reinforcement Learning

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

2026-06-25 · Sicheng Zhang, Muzammal Naseer, Binzhu Xie, Naufal Suryanto 외 arxiv

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applications increasingly demand visually ground…

Continual Pretraining

CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense

2019-08-08 · Difei Gao, Ruiping Wang, Shiguang Shan, Xilin Chen

Alternatively inferring on the visual facts and commonsense is fundamental for an advanced VQA system. This ability requires models to go beyond the literal understanding of commonsense. The system should not just treat …

Question AnsweringVisual Question Answering (VQA)

VisualCOMET: Reasoning about the Dynamic Context of a Still Image

2020-04-22 · ECCV 2020 8 · Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi 외

Even from a single frame of a still image, people can reason about the dynamic story of the image before, after, and beyond the frame. For example, given an image of a man struggling to stay afloat in water, we can reaso…

Visual Commonsense Reasoning