paper-with-me

홈 › Papers

Cognitive Visual Commonsense Reasoning Using Dynamic Working Memory

2021-07-04 · Xuejiao Tang, Xin Huang, Wenbin Zhang, Travers B. Child, Qiong Hu, Zhen Liu, Ji Zhang

Visual Commonsense Reasoning (VCR) predicts an answer with corresponding rationale, given a question-image input. VCR is a recently introduced visual scene understanding task with a wide range of applications, including visual question answering, automated vehicle systems, and clinical decision support. Previous approaches to solving the VCR task generally rely on pre-training or exploiting memory with long dependency relationship encoded models. However, these approaches suffer from a lack of generalizability and prior knowledge. In this paper we propose a dynamic working memory based cognitive VCR network, which stores accumulated commonsense between sentences to provide prior knowledge for inference. Extensive experiments show that the proposed model yields significant improvements over existing methods on the benchmark VCR dataset. Moreover, the proposed model provides intuitive interpretation into visual commonsense reasoning. A Python implementation of our mechanism is publicly available at https://github.com/tanjatang/DMVCR

📄 PDF Abstract BibTeX arXiv:2107.01671

Code (1)

tanjatang/DMVCR 공식 구현 pytorch

Tasks

Question AnsweringScene UnderstandingVisual Commonsense ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Interpretable Visual Understanding with Cognitive Attention Network

2021-08-06 · Xuejiao Tang, Wenbin Zhang, Yi Yu, Kea Turner 외

While image understanding on recognition-level has achieved remarkable advancements, reliable visual scene understanding requires comprehensive image understanding on recognition-level but also cognition-level, which cal…

Scene UnderstandingVisual Commonsense Reasoning

KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning

2020-12-13 · Dandan song, Siyi Ma, Zhanchen Sun, Sicheng Yang 외

Reasoning is a critical ability towards complete visual understanding. To develop machine with cognition-level visual understanding and reasoning abilities, the visual commonsense reasoning (VCR) task has been introduced…

SentenceVisual Commonsense ReasoningVisual Question Answering (VQA)

EventLens: Leveraging Event-Aware Pretraining and Cross-modal Linking Enhances Visual Commonsense Reasoning

2024-04-22 · Mingjie Ma, zhihuan yu, Yichao Ma, GuoHui Li

Visual Commonsense Reasoning (VCR) is a cognitive task, challenging models to answer visual questions requiring human commonsense, and to provide rationales explaining why the answers are correct. With emergence of Large…

Visual Commonsense Reasoning

To Root Artificial Intelligence Deeply in Basic Science for a New Generation of AI

2020-09-11 · Jingan Yang, Yang Peng

One of the ambitions of artificial intelligence is to root artificial intelligence deeply in basic science while developing brain-inspired artificial intelligence platforms that will promote new scientific discoveries. T…

Brain Computer InterfaceDecision MakingVisual Commonsense Reasoning

CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments

2025-10-29 · Rishika Bhagwatkar, Syrielle Montariol, Angelika Romanou, Beatriz Borges 외 arxiv

Humans can naturally identify, reason about, and explain anomalies in their environment. In computer vision, this long-standing challenge remains limited to industrial defects or unrealistic, synthetically generated anom…

Anomaly DetectionVisual Grounding