paper-with-me

홈 › Papers

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

2026-02-28 · Elias Malomgré, Pieter Simoens arxiv

Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain essential as increasingly powerful autonomous decision components are embedded within agent-based systems. While learned and generative models substantially expand system capability, their safety behavior is often entangled with training, making it opaque, difficult to audit, and costly to update after deployment. This paper formalizes the Alignment Flywheel as a governance-centric hybrid MAS architecture that decouples decision generation from safety governance. A Proposer, representing any autonomous decision component, generates candidate trajectories, while a Safety Oracle returns raw safety signals through a stable interface. An enforcement layer applies explicit risk policy at runtime, and a governance MAS supervises the Oracle through auditing, uncertainty-driven verification, and versioned refinement. The central engineering principle is patch locality: many newly observed safety failures can be mitigated by updating the governed oracle artifact and its release pipeline rather than retracting or retraining the underlying decision component. The architecture is implementation-agnostic with respect to both the Proposer and the Safety Oracle, and specifies the roles, artifacts, protocols, and release semantics needed for runtime gating, audit intake, signed patching, and staged rollout across distributed deployments. The result is a hybrid MAS engineering framework for integrating highly capable but fallible autonomous systems under explicit, version-controlled, and auditable oversight.

📄 PDF Abstract BibTeX arXiv:2603.02259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment

2026-02-16 · Elias Malomgré, Pieter Simoens arxiv

AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, dif…

Reinforcement Learning

Optimal Co-Design of a Hybrid Energy Storage System for Truck Charging

2025-06-02 · Juan Pablo Bertucci, Sudarshan Raghuraman, Mauro Salazar, Theo Hofman

The major challenges to battery electric truck adoption are their high cost and grid congestion.In this context, stationary energy storage systems can help mitigate both issues. Since their design and operation are stron…

The Economics of AI Foundation Models: Openness, Competition, and Governance

2025-10-17 · Fasheng Xu, Xiaoyu Wang, Wei Chen, Karen Xie arxiv

The strategic choice of model "openness" has become a defining issue for the foundation model (FM) ecosystem. While this choice is intensely debated, its underlying economic drivers remain underexplored. We construct a t…

PLLuM: A Family of Polish Large Language Models

2025-11-05 · Jan Kocoń, Maciej Piasecki, Arkadiusz Janz, Teddy Ferdinan 외 arxiv

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish …

New Online Communities: Graph Deep Learning on Anonymous Voting Networks to Identify Sybils in Polycentric Governance

2023-11-25 · Quinn DuPont

This research examines the polycentric governance of digital assets in blockchain-based Decentralized Autonomous Organizations (DAOs). It offers a theoretical framework and addresses a critical challenge facing decentral…