paper-with-me

홈 › Papers

The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks

2024-10-09 · Isaac R. Galatzer-Levy, David Munday, Jed McGiffin, Xin Liu, Danny Karmon, Ilia Labzovsky, Rivka Moroshko, Amir Zait, Daniel McDuff

There is increasing interest in tracking the capabilities of general intelligence foundation models. This study benchmarks leading large language models and vision language models against human performance on the Wechsler Adult Intelligence Scale (WAIS-IV), a comprehensive, population-normed assessment of underlying human cognition and intellectual abilities, with a focus on the domains of VerbalComprehension (VCI), Working Memory (WMI), and Perceptual Reasoning (PRI). Most models demonstrated exceptional capabilities in the storage, retrieval, and manipulation of tokens such as arbitrary sequences of letters and numbers, with performance on the Working Memory Index (WMI) greater or equal to the 99.5th percentile when compared to human population normative ability. Performance on the Verbal Comprehension Index (VCI) which measures retrieval of acquired information, and linguistic understanding about the meaning of words and their relationships to each other, also demonstrated consistent performance at or above the 98th percentile. Despite these broad strengths, we observed consistently poor performance on the Perceptual Reasoning Index (PRI; range 0.1-10th percentile) from multimodal models indicating profound inability to interpret and reason on visual information. Smaller and older model versions consistently performed worse, indicating that training data, parameter count and advances in tuning are resulting in significant advances in cognitive ability.

📄 PDF Abstract BibTeX arXiv:2410.07391

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Generative AI as a metacognitive agent: A comparative mixed-method study with human participants on ICF-mimicking exam performance

2024-05-07 · Jelena Pavlovic, Jugoslav Krstic, Luka Mitrovic, Djordje Babic 외

This study investigates the metacognitive capabilities of Large Language Models relative to human metacognition in the context of the International Coaching Federation ICF mimicking exam, a situational judgment test rela…

Relation

Amplifying Human Creativity and Problem Solving with AI Through Generative Collective Intelligence

2025-05-25 · Thomas P. Kehler, Scott E. Page, Alex Pentland, Martin Reeves 외

We propose a new framework for human-AI collaboration that amplifies the distinct capabilities of both. This framework, which we call Generative Collective Intelligence (GCI), shifts AI to the group/social level and empl…

Decision Making

Amplification to Synthesis: A Comparative Analysis of Cognitive Operations Before and After Generative AI

2026-05-13 · Liz Cho, Dongwook Yoon arxiv

Cognitive operations are a rising concern in the geopolitical sphere, a quiet yet rigorous fight for public perception and decision making. While such operations have been extensively studied in the context of bot-driven…

Decision Making

Using profiles of cognitive capability to assess AI suitability for workplace tasks

2026-08-26 · Jonathan Prunty, Marko Tešić, Patrick Quinn, José Hernández-Orallo 외 arxiv

Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where sy…

Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness

2025-11-23 · Svitlana Volkova, Will Dupree, Hsien-Te Kao, Peter Bautista 외 arxiv

This paper introduces BRIES, a novel compound AI architecture designed to detect and measure the effectiveness of persuasion attacks across information environments. We present a system with specialized agents: a Twister…

Prompt EngineeringCausal Inference