paper-with-me

Papers

Possible principles for aligned structure learning agents

2024-09-30 · Lancelot Da Costa, Tomáš Gavenčiak, David Hyland, Mandana Samiei, Cristian Dragos-Manta, Candice Pattisapu, Adeel Razi, Karl Friston

This paper offers a roadmap for the development of scalable aligned artificial intelligence (AI) from first principle descriptions of natural intelligence. In brief, a possible path toward scalable aligned AI rests upon enabling artificial agents to learn a good model of the world that includes a good model of our preferences. For this, the main objective is creating agents that learn to represent the world and other agents' world models; a problem that falls under structure learning (a.k.a. causal representation learning). We expose the structure learning and alignment problems with this goal in mind, as well as principles to guide us forward, synthesizing various ideas across mathematics, statistics, and cognitive science. 1) We discuss the essential role of core knowledge, information geometry and model reduction in structure learning, and suggest core structural modules to learn a wide range of naturalistic worlds. 2) We outline a way toward aligned agents through structure learning and theory of mind. As an illustrative example, we mathematically sketch Asimov's Laws of Robotics, which prescribe agents to act cautiously to minimize the ill-being of other agents. We supplement this example by proposing refined approaches to alignment. These observations may guide the development of artificial intelligence in helping to scale existing -- or design new -- aligned structure learning systems.

📄 PDF Abstract BibTeX arXiv:2410.00258

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

LLM Agents Should Employ Security Principles

2025-05-29 · Kaiyuan Zhang, Zian Su, Pin-Yu Chen, Elisa Bertino 외

Large Language Model (LLM) agents show considerable promise for automating complex tasks using contextual reasoning; however, interactions involving multiple agents and the system's susceptibility to prompt injection and…

Large Language Model

Social, Legal, Ethical, Empathetic and Cultural Norm Operationalisation for AI Agents

2026-03-12 · Radu Calinescu, Ana Cavalcanti, Marsha Chechik, Lina Marsso 외 arxiv

As AI agents are increasingly used in high-stakes domains like healthcare and law enforcement, aligning their behaviour with social, legal, ethical, empathetic, and cultural (SLEEC) norms has become a critical engineerin…

Towards Self-constructive Artificial Intelligence: Algorithmic basis (Part I)

2019-01-06 · Fernando J. Corbacho

Artificial Intelligence frameworks should allow for ever more autonomous and general systems in contrast to very narrow and restricted (human pre-defined) domain systems, in analogy to how the brain works. Self-construct…

A Selfish Herd with a Target

2024-10-17 · Thomas Stemler, Shannon Dee Algar, Jesse Zhou

One of the most striking phenomena in biological systems is the tendency for biological agents to spatially aggregate, and subsequently display further collective behaviours such as rotational motion. One prominent expla…

Towards Rigorous Design of OoD Detectors

2023-06-14 · Chih-Hong Cheng, Changshun Wu, Harald Ruess, Saddek Bensalem

Out-of-distribution (OoD) detection techniques are instrumental for safety-related neural networks. We are arguing, however, that current performance-oriented OoD detection techniques geared towards matching metrics such…

Out of Distribution (OOD) Detection