Visual Reasoning
12개 벤치마크 · 논문 1,326편 · 이 태스크의 논문 보기 →
Benchmarks
Winoground
NLVR2 Dev
NLVR2 Test
CLEVRER
Bongard-OpenWorld
WinoGAViL
VSR
PHYRE-1B-Cross
PHYRE-1B-Within
VASR
NLVR
Most implemented
Learning Transferable Visual Models From Natural Language Supervision
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
Visual Instruction Tuning
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
VisualBERT: A Simple and Performant Baseline for Vision and Language
Compositional Attention Networks for Machine Reasoning
Papers
Reason Through the Latent! Making Latent Visual Reasoning Necessary
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply tha…
Multimodal ReasoningVisual ReasoningFrom Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-bas…
Visual ReasoningPACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two f…
Visual ReasoningG2D: Generative-to-Discriminative Collaborative Inference for Zero-Shot Image Classification
Zero-shot classification needs efficient label retrieval and fine-grained visual reasoning, yet discriminative and generative vision-language models fail in complementary ways.When CLIP's top-1 prediction is wrong, the c…
Zero-Shot Image ClassificationVisual ReasoningVBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for …
Reinforcement LearningVideo GenerationVisual ReasoningInvestigating Relational Reasoning in VLMs
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such as language cues or priors. To investig…
Relational ReasoningVisual Reasoning