paper-with-me

홈 › Papers

The Catastrophic Paradox of Human Cognitive Frameworks in Large Language Model Evaluation: A Comprehensive Empirical Analysis of the CHC-LLM Incompatibility

2025-11-23 · Mohan Reddy arxiv

This investigation presents an empirical analysis of the incompatibility between human psychometric frameworks and Large Language Model evaluation. Through systematic assessment of nine frontier models including GPT-5, Claude Opus 4.1, and Gemini 3 Pro Preview using the Cattell-Horn-Carroll theory of intelligence, we identify a paradox that challenges the foundations of cross-substrate cognitive evaluation. Our results show that models achieving above-average human IQ scores ranging from 85.0 to 121.4 simultaneously exhibit binary accuracy rates approaching zero on crystallized knowledge tasks, with an overall judge-binary correlation of r = 0.175 (p = 0.001, n = 1800). This disconnect appears most strongly in the crystallized intelligence domain, where every evaluated model achieved perfect binary accuracy while judge scores ranged from 25 to 62 percent, which cannot occur under valid measurement conditions. Using statistical analyses including Item Response Theory modeling, cross-vendor judge validation, and paradox severity indexing, we argue that this disconnect reflects a category error in applying biological cognitive architectures to transformer-based systems. The implications extend beyond methodology to challenge assumptions about intelligence, measurement, and anthropomorphic biases in AI evaluation. We propose a framework for developing native machine cognition assessments that recognize the non-human nature of artificial intelligence.

📄 PDF Abstract BibTeX arXiv:2511.18302

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Augmentation to Symbiosis: A Review of Human-AI Collaboration Frameworks, Performance, and Perils

2025-11-07 · Richard Jiarui Tong arxiv

This paper offers a concise, 60-year synthesis of human-AI collaboration, from Licklider's ``man-computer symbiosis" (AI as colleague) and Engelbart's ``augmenting human intellect" (AI as tool) to contemporary poles: Hum…

Zero-Forgetting CISS via Dual-Phase Cognitive Cascades

2026-03-14 · Yuquan Lu, Yifu Guo, Zishan Xu, Siyu Zhang 외 arxiv

Continual semantic segmentation (CSS) is a cornerstone task in computer vision that enables a large number of downstream applications, but faces the catastrophic forgetting challenge. In conventional class-incremental se…

Continual Semantic SegmentationContinual Learning

Unifying Decision-Making: a Review on Evolutionary Theories on Rationality and Cognitive Biases

2018-11-29 · Catarina Moreira

In this paper, we make a review on the concepts of rationality across several different fields, namely in economics, psychology and evolutionary biology and behavioural ecology. We review how processes like natural selec…

Decision Making

Resolution of the St. Petersburg paradox using Von Mises axiom of randomness

2019-06-27

In this article we will propose a completely new point of view for solving one of the most important paradoxes concerning game theory. The solution develop shifts the focus from the result to the strategy s ability to op…

Decision Making

Quantum-like Structure in Multidimensional Relevance Judgements

2020-01-20 · Sagar Uprety, Prayag Tiwari, Shahram Dehdashti, Lauren Fell 외

A large number of studies in cognitive science have revealed that probabilistic outcomes of certain human decisions do not agree with the axioms of classical probability theory. The field of Quantum Cognition provides an…

Decision MakingDecision Making Under Uncertainty