paper-with-me

홈 › Papers

SuperARC: An Agnostic Test for Narrow, General, and Super Intelligence Based On the Principles of Recursive Compression and Algorithmic Probability

2025-03-20 · Alberto Hernández-Espinosa, Luan Ozelim, Felipe S. Abrahão, Hector Zenil

We introduce an open-ended test grounded in algorithmic probability that can avoid benchmark contamination in the quantitative evaluation of frontier models in the context of their Artificial General Intelligence (AGI) and Superintelligence (ASI) claims. Unlike other tests, this test does not rely on statistical compression methods (such as GZIP or LZW), which are more closely related to Shannon entropy than to Kolmogorov complexity and are not able to test beyond simple pattern matching. The test challenges aspects of AI, in particular LLMs, related to features of intelligence of fundamental nature such as synthesis and model creation in the context of inverse problems (generating new knowledge from observation). We argue that metrics based on model abstraction and abduction (optimal Bayesian inference') for predictive planning' can provide a robust framework for testing intelligence, including natural intelligence (human and animal), narrow AI, AGI, and ASI. We found that LLM model versions tend to be fragile and incremental as a result of memorisation only with progress likely driven by the size of training data. The results were compared with a hybrid neurosymbolic approach that theoretically guarantees universal intelligence based on the principles of algorithmic probability and Kolmogorov complexity. The method outperforms LLMs in a proof-of-concept on short binary sequences. We prove that compression is equivalent and directly proportional to a system's predictive power and vice versa. That is, if a system can better predict it can better compress, and if it can better compress, then it can better predict. Our findings strengthen the suspicion regarding the fundamental limitations of LLMs, exposing them as systems optimised for the perception of mastery over human language.

📄 PDF Abstract BibTeX arXiv:2503.16743

Code (1)

AlgoDynLab/SuperintelligenceTest 공식 구현 pytorch

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs

2026-01-05 · Jinghan Ru, Siyuan Yan, Yuguo Yin, Yuexian Zou 외 arxiv

Multimodal Large Language Models (MLLMs) show promise for medical applications, yet progress in dermatology lags due to limited training data, narrow task coverage, and lack of clinically-grounded supervision that mirror…

Reinforcement LearningTest-time Adaptation

Emergent alignment and the projectability of ethical personas

2026-06-08 · Guillermo Del Pinal, Youngchan Lee, Calum McNamara, Alejandro Perez Carballo arxiv

Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate diffe…

A Framework for Searching for General Artificial Intelligence

2016-11-02 · Marek Rosa, Jan Feyereisl, The GoodAI Collective

There is a significant lack of unified approaches to building generally intelligent machines. The majority of current artificial intelligence research operates within a very narrow field of focus, frequently without cons…

Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement

2021-07-27 · Siyuan Zhang, Xiaofei Li

This paper addresses the problem of microphone array generalization for deep-learning-based end-to-end multichannel speech enhancement. We aim to train a unique deep neural network (DNN) potentially performing well on un…

Speech Enhancement

Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness

2026-08-10 · Tadanobu Chuyo Kamijo, Ori Rottenstreich, Javier Conde, Gonzalo Martínez 외 hf

Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world dep…