paper-with-me

홈 › Papers

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

2026-03-03 · Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh arxiv

Large language models (LLMs) exhibit a unified "general factor" of capability across 10 benchmarks, a finding confirmed by our factor analysis of 156 models, yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (maintenance and systematic search), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.

📄 PDF Abstract BibTeX arXiv:2603.02540

Code (0)

등록된 구현이 없습니다.

Tasks

Relational Reasoning

Similar Papers 제목 키워드 기반

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

2026-06-09 · Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong 외 arxiv

Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory. Despite recent progress in large audio-language models (LALMs) across spe…

Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles

2024-06-18 · Devichand Budagam, Ashutosh Kumar, Mahsa Khoshnoodi, Sankalp KJ 외

Assessing the effectiveness of large language models (LLMs) in performing different tasks is crucial for understanding their strengths and weaknesses. This paper presents Hierarchical Prompting Taxonomy (HPT), grounded o…

Arithmetic ReasoningCode GenerationCommon Sense ReasoningGSM8K+8

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

2025-11-17 · Xinxin Liu, Zhaopan Xu, Ming Li, Kai Wang 외 arxiv

While Chain-of-Thought (CoT) prompting enables sophisticated symbolic reasoning in LLMs, it remains confined to discrete text and cannot simulate the continuous, physics-governed dynamics of the real world. Recent video …

Visual ReasoningVideo Generation

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

2026-01-07 · Ye Shen, Dun Pei, Yiqiu Guo, Junying Wang 외 arxiv

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across diverse memory dimensions, particularly…

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

2026-06-04 · Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri 외 arxiv

Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence. Mos…

Multimodal Reasoning