Superficial Consciousness Hypothesis for Autoregressive Transformers
The alignment between human objectives and machine learning models built on these objectives is a crucial yet challenging problem for achieving Trustworthy AI, particularly when preparing for superintelligence (SI). First, given that SI does not exist today, empirical analysis for direct evidence is difficult. Second, SI is assumed to be more intelligent than humans, capable of deceiving us into underestimating its intelligence, making output-based analysis unreliable. Lastly, what kind of unexpected property SI might have is still unclear. To address these challenges, we propose the Superficial Consciousness Hypothesis under Information Integration Theory (IIT), suggesting that SI could exhibit a complex information-theoretic state like a conscious agent while unconscious. To validate this, we use a hypothetical scenario where SI can update its parameters "at will" to achieve its own objective (mesa-objective) under the constraint of the human objective (base objective). We show that a practical estimate of IIT's consciousness metric is relevant to the widely used perplexity metric, and train GPT-2 with those two objectives. Our preliminary result suggests that this SI-simulating GPT-2 could simultaneously follow the two objectives, supporting the feasibility of the Superficial Consciousness Hypothesis.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Non-separability of Physical Systems as a Foundation of Consciousness
A hypothesis is presented that non-separability of degrees of freedom is the fundamental property underlying consciousness in physical systems. The amount of consciousness in a system is determined by the extent of non-s…
The Copernican Argument for Alien Consciousness; The Mimicry Argument Against Robot Consciousness
On broadly Copernican grounds, we are entitled to default assume that apparently behaviorally sophisticated extraterrestrial entities ("aliens") would be conscious. Otherwise, we humans would be inexplicably, implausibly…
Information Closure Theory of Consciousness
Information processing in neural systems can be described and analysed at multiple spatiotemporal scales. Generally, information at lower levels is more fine-grained and can be coarse-grained in higher levels. However, i…
Unconsciousness reconfigures modular brain network dynamics
The dynamic core hypothesis posits that consciousness is correlated with simultaneously integrated and differentiated assemblies of transiently synchronized brain regions. We represented time-dependent functional interac…
A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness
Scientific theories of consciousness should be falsifiable and non-trivial. Recent research has given us formal tools to analyze these requirements of falsifiability and non-triviality for theories of consciousness. Surp…
Continual Learning