paper-with-me

Papers

The Reasons that Agents Act: Intention and Instrumental Goals

2024-02-11 · Francis Rhys Ward, Matt MacDermott, Francesco Belardinelli, Francesca Toni, Tom Everitt

Intention is an important and challenging concept in AI. It is important because it underlies many other concepts we care about, such as agency, manipulation, legal responsibility, and blame. However, ascribing intent to AI systems is contentious, and there is no universally accepted theory of intention applicable to AI agents. We operationalise the intention with which an agent acts, relating to the reasons it chooses its decision. We introduce a formal definition of intention in structural causal influence models, grounded in the philosophy literature on intent and applicable to real-world machine learning systems. Through a number of examples and results, we show that our definition captures the intuitive notion of intent and satisfies desiderata set-out by past work. In addition, we show how our definition relates to past concepts, including actual causality, and the notion of instrumental goals, which is a core idea in the literature on safe AI agents. Finally, we demonstrate how our definition can be used to infer the intentions of reinforcement learning agents and language models from their behaviour.

📄 PDF Abstract BibTeX arXiv:2402.07221

Code (0)

등록된 구현이 없습니다.

Tasks

Philosophy

Similar Papers 제목 키워드 기반

An Explainable Collaborative Dialogue System using a Theory of Mind

2023-02-19 · Philip R. Cohen, Lucian Galescu, Maayan Shvo

Eva is a neuro-symbolic domain-independent multimodal collaborative dialogue system that takes seriously that the purpose of task-oriented dialogue is to assist the user. To do this, the system collaborates by inferring …

Argumentation-based Agents that Explain their Decisions

2020-09-13 · Mariela Morveli-Espinoza, Ayslan Possebom, Cesar Augusto Tacla

Explainable Artificial Intelligence (XAI) systems, including intelligent agents, must be able to explain their internal decisions, behaviours and reasoning that produce their choices to the humans (or other systems) with…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?

2025-02-16 · Yufei He, Yuexin Li, Jiaying Wu, Yuan Sui 외

As large language models (LLMs) continue to evolve, ensuring their alignment with human goals and values remains a pressing challenge. A key concern is \textit{instrumental convergence}, where an AI system, in optimizing…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Towards Using Promises for Multi-Agent Cooperation in Goal Reasoning

2022-06-20 · Daniel Swoboda, Till Hofmann, Tarik Viehmann, Gerhard Lakemeyer

Reasoning and planning for mobile robots is a challenging problem, as the world evolves over time and thus the robot's goals may change. One technique to tackle this problem is goal reasoning, where the agent not only re…

Will artificial agents pursue power by default?

2025-06-02 · Christian Tarsney

Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that…