Visual Commonsense Reasoning
7개 벤치마크 · 논문 68편 · 이 태스크의 논문 보기 →
Benchmarks
GD-VCR
VCR (Q-A) dev
VCR (Q-A) test
VCR (Q-AR) dev
VCR (Q-AR) test
VCR (QA-R) dev
VCR (QA-R) test
Most implemented
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
UNITER: UNiversal Image-TExt Representation Learning
From Recognition to Cognition: Visual Commonsense Reasoning
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Papers
Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation
Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture reliable visual commonsense knowledge. We propose a model-agnostic, s…
Visual Commonsense ReasoningScene Graph GenerationCausal Debiasing for Visual Commonsense Reasoning
Visual Commonsense Reasoning (VCR) refers to answering questions and providing explanations based on images. While existing methods achieve high prediction accuracy, they often overlook bias in datasets and lack debiasin…
Visual Commonsense ReasoningThe Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation from Recognition to Reasoning
This survey paper chronicles the evolution of evaluation in multimodal artificial intelligence (AI), framing it as a progression of increasingly sophisticated "cognitive examinations." We argue that the field is undergoi…
Visual Commonsense ReasoningCompositional Image-Text Matching and Retrieval by Grounding Entities
Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these models excel in various downstream tasks, inc…
Image CaptioningImage-text matchingQuestion AnsweringRetrieval+3Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
Visual Commonsense Reasoning, which is regarded as one challenging task to pursue advanced visual scene comprehension, has been used to diagnose the reasoning ability of AI systems. However, reliable reasoning requires a…
Visual Commonsense ReasoningHow Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey
The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the rese…
Image CaptioningQuestion AnsweringVisual Commonsense ReasoningVisual Question Answering