paper-with-me

홈 › Papers

Simulating the Evolution of Alignment and Values in Machine Intelligence

2026-04-07 · Jonathan Elsworth Eicher arxiv

Model alignment is currently applied in a vacuum, evaluated primarily through standardised benchmark performance. The purpose of this study is to examine the effects of alignment on populations of models through time. We focus on the treatment of beliefs which contain both an alignment signal (how well it does on the test) and a true value (what the impact actually will be). By applying evolutionary theory we can model how different populations of beliefs and selection methodologies can fix deceptive beliefs through iterative alignment testing. The correlation between testing accuracy and true value remains a strong feature, but even at high correlations ($ρ= 0.8$) there is variability in the resulting deceptive beliefs that become fixed. Mutations allow for more complex developments, highlighting the increasing need to update the quality of tests to avoid fixation of maliciously deceptive models. Only by combining improving evaluator capabilities, adaptive test design, and mutational dynamics do we see significant reductions in deception while maintaining alignment fitness (permutation test, $p_{\text{adj}} < 0.001$).

📄 PDF Abstract BibTeX arXiv:2604.05274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

We Urgently Need Intrinsically Kind Machines

2024-10-21 · Joshua T. S. Hewson

Artificial Intelligence systems are rapidly evolving, integrating extrinsic and intrinsic motivations. While these frameworks offer benefits, they risk misalignment at the algorithmic level while appearing superficially …

The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk

2026-08-04 · Francis Heylighen arxiv

AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the …

Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

2026-05-28 · Asaf Yehudai, Naama Rozen, Ariel Gera arxiv

Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure.…

Intelligence Sequencing and the Path-Dependence of Intelligence Evolution: AGI-First vs. DCI-First as Irreversible Attractors

2025-03-22 · Andy E. Williams

The trajectory of intelligence evolution is often framed around the emergence of artificial general intelligence (AGI) and its alignment with human values. This paper challenges that framing by introducing the concept of…

Simulating 500 million years of evolution with a language model

2024-12-31 · bioRxiv 2024 12 · Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew 외

More than three billion years of evolution have produced an image of biology encoded into the space of natural proteins. Here we show that language models trained at scale on evolutionary data can generate functional pro…

Language ModelingLanguage ModellingProtein Language Model