paper-with-me

홈 › Papers

From Human-Level AI Tales to AI Leveling Human Scales

2026-02-21 · Peter Romero, Fernando Martínez-Plumed, Zachary R. Tidler, Matthieu Téhénan, Sipeng Chen, Álvaro David Gómez Antón, Luning Sun, Manuel Cebrian, Lexin Zhou, Yael Moros Daval, Daniel Romero-Alvarado, Félix Martí Pérez, Kevin Wei, José Hernández-Orallo arxiv

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world population' and report performance on a common, human-anchored scale. Concretely, we build on a set of multi-level scales for different capabilities where each level should represent a probability of success of the whole world population on a logarithmic scale with a base $B$. We calibrate each scale for each capability (reasoning, comprehension, knowledge, volume, etc.) by compiling publicly released human test data spanning education and reasoning benchmarks (PISA, TIMSS, ICAR, UKBioBank, and ReliabilityBench). The base $B$ is estimated by extrapolating between samples with two demographic profiles using LLMs, with the hypothesis that they condense rich information about human populations. We evaluate the quality of different mappings using group slicing and post-stratification. The new techniques allow for the recalibration and standardization of scales relative to the whole-world population.

📄 PDF Abstract BibTeX arXiv:2602.18911

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human-AI Collaborative Bot Detection in MMORPGs

2025-08-28 · Jaeman Son, Hyunsoo Kim arxiv

In Massively Multiplayer Online Role-Playing Games (MMORPGs), auto-leveling bots exploit automated programs to level up characters at scale, undermining gameplay balance and fairness. Detecting such bots is challenging, …

Representation Learning

Leveling3D: Leveling Up 3D Reconstruction with Feed-Forward 3D Gaussian Splatting and Geometry-Aware Generation

2026-03-17 · Yiming Huang, Baixiang Huang, Beilei Cui, Chi Kit Ng 외 arxiv

Feed-forward 3D reconstruction has revolutionized 3D vision, providing a powerful baseline for downstream tasks such as novel-view synthesis with 3D Gaussian Splatting. Previous works explore fixing the corrupted renderi…

3D ReconstructionDepth Estimation

Talking to Extraordinary Objects: Folktales Offer Analogies for Interacting with Technology

2026-01-10 · Martha Larson arxiv

Speech and language are valuable for interacting with technology. It would be ideal to be able to decouple their use from anthropomorphization, which has recently met an important moment of reckoning. In the world of fol…

Not All Features Are Equal: Feature Leveling Deep Neural Networks for Better Interpretation

2019-05-24 · ICLR 2020 1 · Yingjing Lu, Runde Yang

Self-explaining models are models that reveal decision making parameters in an interpretable manner so that the model reasoning process can be directly understood by human beings. General Linear Models (GLMs) are self-ex…

AllDecision Making

TALES: Text Adventure Learning Environment Suite

2025-04-19 · Christopher Zhang Cui, Xingdi Yuan, Zhang Xiao, Prithviraj Ammanabrolu 외

Reasoning is an essential skill to enable Large Language Models (LLMs) to interact with the world. As tasks become more complex, they demand increasingly sophisticated and diverse reasoning capabilities for sequential de…

Decision MakingSequential Decision Making