paper-with-me

홈 › Papers

The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension

2026-06-26 · Jose Pascual Gumbau Mezquita arxiv

We establish the mathematical limits of AGI safety in two forms: verifying a fixed system, and verifying that a certified safety property persists once the system self-modifies. In the static case, no algorithm can certify a highly expressive AGI's safe behaviour infallibly, completely and tractably, whether over unbounded input domains (blocked by Rice's and Godel's theorems) or over all finite hardware configurations (blocked by Trakhtenbrot's theorem, which splits into a PSPACE-hardness barrier and a co-RE-completeness barrier), forcing a Soundness-Completeness-Tractability Trilemma as a structural, not statistical, necessity. In the dynamic case, we formalise self-modification as a computable transition operator and prove that no algorithm can determine, from a system's current certified safety, whether safety survives its next self-modification step: a result that reduces to Rice's Theorem one level up, making the static and dynamic barriers two faces of one obstruction. This forces an exclusive dichotomy: persistent certification is attainable only for systems that have stopped evolving semantically, i.e. only for narrow, not general, systems. Nor can the obstruction be delegated: any supervisor adequate to audit a general AGI is itself a general AGI, so the supervisory regress never terminates. Three practical risks (finite test coverage, bounded deliberation time, restricted observation) are one phenomenon: every bounded scheme that does not reject correct evidence admits an evolution trace it certifies at every stage while the property is persistently violated. These results give formal content to the unverifiability of AI, showing it is not an engineering target deferred by current limits but a structural tension, an Expressivity Invariant governed by the same computational laws as the Halting Problem and Rice's Theorem.

📄 PDF Abstract BibTeX arXiv:2606.28639

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Verifier Theory and Unverifiability

2016-09-01 · Roman V. Yampolskiy

Despite significant developments in Proof Theory, surprisingly little attention has been devoted to the concept of proof verifier. In particular, the mathematical community may be interested in studying different types o…

Automated Theorem ProvingGeneral Classification

Super Co-alignment for Sustainable Symbiotic Society

2025-04-24 · Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu 외

As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead …

Asymptotically Unambitious Artificial General Intelligence

2019-05-29 · Michael K. Cohen, Badri Vellambi, Marcus Hutter

General intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen imp…

Self-Driving Cars

A General Theory of Growth, Employment, and Technological Change: Experiential Matrix Theory and the Transition from GDP to Humanist Experiential Growth in the Age of Artificial Intelligence

2025-05-25 · Christian Callaghan

This paper introduces Experiential Matrix Theory (EMT), a general theory of growth, employment, and technological change for the age of artificial intelligence (AI). EMT redefines utility as the alignment between product…

Implications of Quantum Computing for Artificial Intelligence alignment research

2019-08-19 · Jaime Sevilla, Pablo Moreno

We explain some key features of quantum computing via three heuristics and apply them to argue that a deep understanding of quantum computing is unlikely to be helpful to address current bottlenecks in Artificial Intelli…