paper-with-me

Papers

Measuring Intelligence Beyond Human Scale

2026-07-08 · Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal, Andrew Tu, Kia Ghods, Mark Braverman, Elad Hazan arxiv

How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.

📄 PDF Abstract BibTeX arXiv:2607.07040

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI or Human? Understanding Perceptions of Embodied Robots with LLMs

2025-07-22 · Lavinia Hriscu, Alberto Sanfeliu, Anais Garrell arxiv

The pursuit of artificial intelligence has long been associated to the the challenge of effectively measuring intelligence. Even if the Turing Test was introduced as a means of assessing a system intelligence, its releva…

Information Retrieval

Measuring Machine Intelligence Through Visual Question Answering

2016-08-31 · C. Lawrence Zitnick, Aishwarya Agrawal, Stanislaw Antol, Margaret Mitchell 외

As machines have become more intelligent, there has been a renewed interest in methods for measuring their intelligence. A common approach is to propose tasks for which a human excels, but one which machines find difficu…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Beyond accuracy: quantifying trial-by-trial behaviour of CNNs and humans by measuring error consistency

2020-06-30 · NeurIPS 2020 12 · Robert Geirhos, Kristof Meding, Felix A. Wichmann

A central problem in cognitive science and behavioural neuroscience as well as in machine learning and artificial intelligence research is to ascertain whether two or more decision makers (be they brains or algorithms) u…

Decision MakingObject Recognition

A Measure for Level of Autonomy Based on Observable System Behavior

2024-07-20 · Jason M. Pittman

Contemporary artificial intelligence systems are pivotal in enhancing human efficiency and safety across various domains. One such domain is autonomous systems, especially in automotive and defense use cases. Artificial …

Decision Making

On the Measure of Intelligence

2019-11-05 · François Chollet

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that en…

ARCBenchmarking