paper-with-me

Papers

Measuring Progress Toward AGI: A Cognitive Framework

2026-05-27 · Ryan Burnell, Yumeya Yamamori, Orhan Firat, Kate Olszewska, Steph Hughes-Fitt, Oran Kelly, Isaac R. Galatzer-Levy, Meredith Ringel Morris, Allan Dafoe, Alison M. Snyder, Noah D. Goodman, Matthew Botvinick, Shane Legg arxiv

Despite widespread discussion of AGI, there is no clear framework for measuring progress toward it. This ambiguity fuels subjective claims, makes it difficult to track progress, and risks hindering responsible governance. As a starting point to address this gap, we present a framework for understanding system capabilities in relation to human cognitive abilities. Drawing from decades of research in psychology, neuroscience, and cognitive science, we introduce a Cognitive Taxonomy that deconstructs general intelligence into 10 key cognitive faculties. We then propose a rigorous evaluation protocol in which a system's performance is measured across a suite of targeted, held-out cognitive tasks, generating a 'cognitive profile' that can be used to understand a system's strengths and weaknesses. We hope this framework will provide a practical roadmap and an initial step toward more rigorous, empirical evaluation of AGI.

📄 PDF Abstract BibTeX arXiv:2605.28405

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aletheia: Quantifying Cognitive Conviction in Reasoning Models via Regularized Inverse Confusion Matrix

2026-01-04 · Fanzhe Fu arxiv

In the progressive journey toward Artificial General Intelligence (AGI), current evaluation paradigms face an epistemological crisis. Static benchmarks measure knowledge breadth but fail to quantify the depth of belief. …

Measuring algorithmic interpretability: A human-learning-based framework and the corresponding cognitive complexity score

2022-05-20 · John P. Lalor, Hong Guo

Algorithmic interpretability is necessary to build trust, ensure fairness, and track accountability. However, there is no existing formal measurement method for algorithmic interpretability. In this work, we build upon p…

Fairness

EMP-EVAL: A Framework for Measuring Empathy in Open Domain Dialogues

2023-01-29 · Bushra Amjad, Muhammad Zeeshan, Mirza Omer Beg

Measuring empathy in conversation can be challenging, as empathy is a complex and multifaceted psychological construct that involves both cognitive and emotional components. Human evaluations can be subjective, leading t…

From Consumption to Collaboration: Measuring Interaction Patterns to Augment Human Cognition in Open-Ended Tasks

2025-04-03 · Joshua Holstein, Moritz Diener, Philipp Spitzer

The rise of Generative AI, and Large Language Models (LLMs) in particular, is fundamentally changing cognitive processes in knowledge work, raising critical questions about their impact on human reasoning and problem-sol…

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

2026-05-23 · Ali Şenol, Garima Agrawal, Huan Liu arxiv

Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored to final-answer correctness, providing limited insight into how models reason, how reliably they behave under contextu…