paper-with-me

Papers

What is the Visual Cognition Gap between Humans and Multimodal LLMs?

2024-06-14 · Xu Cao, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Jianguo Cao, James M. Rehg

Recently, Multimodal Large Language Models (MLLMs) have shown great promise in language-guided perceptual tasks such as recognition, segmentation, and object detection. However, their effectiveness in addressing visual cognition problems that require high-level reasoning is not well-established. One such challenge is abstract visual reasoning (AVR) -- the cognitive ability to discern relationships among patterns in a set of images and extrapolate to predict subsequent patterns. This skill is crucial during the early neurodevelopmental stages of children. Inspired by the AVR tasks in Raven's Progressive Matrices (RPM) and Wechsler Intelligence Scale for Children (WISC), we propose a new dataset MaRs-VQA and a new benchmark VCog-Bench containing three datasets to evaluate the zero-shot AVR capability of MLLMs and compare their performance with existing human intelligent investigation. Our comparative experiments with different open-source and closed-source MLLMs on the VCog-Bench revealed a gap between MLLMs and human intelligence, highlighting the visual cognitive limitations of current MLLMs. We believe that the public release of VCog-Bench, consisting of MaRs-VQA, and the inference pipeline will drive progress toward the next generation of MLLMs with human-like visual cognition abilities.

📄 PDF Abstract BibTeX arXiv:2406.10424

Code (1)

IrohXu/VCog-Bench 공식 구현 pytorch

Tasks

object-detectionObject DetectionVisual Question Answering (VQA)Visual Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models

2024-07-25 · Tessa Verhoef, Kiana Shahrasbi, Tom Kouwenhoven

Humans have clear cross-modal preferences when matching certain novel words to visual shapes. Evidence suggests that these preferences play a prominent role in our linguistic processing, language learning, and the origin…

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

2026-08-30 · Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang 외 arxiv

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal l…

Multimodal Reasoning

MERLOT: Multimodal Neural Script Knowledge Models

2021-06-04 · NeurIPS 2021 12 · Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 외

As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future. We introduce MERLOT, a model that learns multimodal sc…

Multimodal ReasoningVisual Commonsense Reasoning

Analyzing Utility of Visual Context in Multimodal Speech Recognition Under Noisy Conditions

2019-06-30 · Tejas Srinivasan, Ramon Sanabria, Florian Metze

Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world. However, it is currently unclear to what extent auxiliary modalities i…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests

2025-10-15 · Fitim Abdullahu, Helmut Grabner arxiv

Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on …