paper-with-me

Papers

Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective

2025-12-04 · Jae Hee Lee, Anne Lauscher, Stefano V. Albrecht arxiv

Large language models (LLMs) have been widely deployed in various applications, often functioning as autonomous agents that interact with each other in multi-agent systems. While these systems have shown promise in enhancing capabilities and enabling complex tasks, they also pose significant ethical challenges. This position paper outlines a research agenda aimed at ensuring the ethical behavior of multi-agent systems of LLMs (MALMs) from the perspective of mechanistic interpretability. We identify three key research challenges: (i) developing comprehensive evaluation frameworks to assess ethical behavior at individual, interactional, and systemic levels; (ii) elucidating the internal mechanisms that give rise to emergent behaviors through mechanistic interpretability; and (iii) implementing targeted parameter-efficient alignment techniques to steer MALMs towards ethical behaviors without compromising their performance.

📄 PDF Abstract BibTeX arXiv:2512.04691

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LOKA Protocol: A Decentralized Framework for Trustworthy and Ethical AI Agent Ecosystems

2025-04-15 · Rajesh Ranjan, Shailja Gupta, Surya Narayan Singh

The rise of autonomous AI agents, capable of perceiving, reasoning, and acting independently, signals a profound shift in how digital ecosystems operate, govern, and evolve. As these agents proliferate beyond centralized…

AI AgentEthics

Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm

2025-09-15 · Alireza Mohamadi, Ali Yavari arxiv

When survival instincts conflict with human welfare, how do Large Language Models (LLMs) make ethical choices? This fundamental tension becomes critical as LLMs integrate into autonomous systems with real-world consequen…

When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems

2026-01-31 · Naen Xu, Hengyu An, Shuo Shi, Jinghuai Zhang 외 arxiv

Recent advancements in large language models (LLMs) have significantly enhanced the capabilities of collaborative multi-agent systems, enabling them to address complex challenges. However, within these multi-agent system…

Mirror: A Multi-Agent System for AI-Assisted Ethics Review

2026-02-09 · Yifan Ding, Yuhui Shi, Zhiyan Li, Zilong Wang 외 arxiv

Ethics review is a foundational mechanism of modern research governance, yet contemporary systems face increasing strain as ethical risks arise as structural consequences of large-scale, interdisciplinary scientific prac…

NAEL: Non-Anthropocentric Ethical Logic

2025-10-16 · Bianca Maria Lerma, Rafael Peñaloza arxiv

We introduce NAEL (Non-Anthropocentric Ethical Logic), a novel ethical framework for artificial agents grounded in active inference and symbolic reasoning. Departing from conventional, human-centred approaches to AI ethi…