paper-with-me

Papers

Zero-shot visual reasoning through probabilistic analogical mapping

2022-09-29 · Taylor W. Webb, Shuhao Fu, Trevor Bihl, Keith J. Holyoak, Hongjing Lu

Human reasoning is grounded in an ability to identify highly abstract commonalities governing superficially dissimilar visual inputs. Recent efforts to develop algorithms with this capacity have largely focused on approaches that require extensive direct training on visual reasoning tasks, and yield limited generalization to problems with novel content. In contrast, a long tradition of research in cognitive science has focused on elucidating the computational principles underlying human analogical reasoning; however, this work has generally relied on manually constructed representations. Here we present visiPAM (visual Probabilistic Analogical Mapping), a model of visual reasoning that synthesizes these two approaches. VisiPAM employs learned representations derived directly from naturalistic visual inputs, coupled with a similarity-based mapping operation derived from cognitive theories of human reasoning. We show that without any direct training, visiPAM outperforms a state-of-the-art deep learning model on an analogical mapping task. In addition, visiPAM closely matches the pattern of human performance on a novel task involving mapping of 3D objects across disparate categories.

📄 PDF Abstract BibTeX arXiv:2209.15087

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis

2024-08-27 · Aishik Nagar, Shantanu Jaiswal, Cheston Tan

Vision-language models (VLMs) have shown impressive zero- and few-shot performance on real-world visual question answering (VQA) benchmarks, alluding to their capabilities as visual reasoning engines. However, the benchm…

BenchmarkingLarge Language ModelQuestion AnsweringVisual Question Answering+3

FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis

2025-01-17 · Zhe Chen, Zijing Chen

Object Referring Analysis (ORA), commonly known as referring expression comprehension, requires the identification and localization of specific objects in an image based on natural descriptions. Unlike generic object det…

Bayesian InferenceLanguage ModelingLanguage ModellingObject+4

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks

2026-03-12 · Mei Chee Leong, Ying Gu, Hui Li Tan, Liyuan Li 외 arxiv

Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, they are often deployed as zero-shot solution to new tasks in a black-box manner. V…

Relational ReasoningLogical Reasoning

Enhancing Zero-shot Commonsense Reasoning by Integrating Visual Knowledge via Machine Imagination

2026-03-05 · Hyuntae Park, Yeachan Kim, SangKeun Lee arxiv

Recent advancements in zero-shot commonsense reasoning have empowered Pre-trained Language Models (PLMs) to acquire extensive commonsense knowledge without requiring task-specific fine-tuning. Despite this progress, thes…

Neural Representational Consistency Emerges from Probabilistic Neural-Behavioral Representation Alignment

2025-05-07 · Yu Zhu, Chunfeng Song, Wanli Ouyang, Shan Yu 외

Individual brains exhibit striking structural and physiological heterogeneity, yet neural circuits can generate remarkably consistent functional properties across individuals, an apparent paradox in neuroscience. While r…