Aligned with Whom? Direct and social goals for AI systems
As artificial intelligence (AI) becomes more powerful and widespread, the AI alignment problem - how to ensure that AI systems pursue the goals that we want them to pursue - has garnered growing attention. This article distinguishes two types of alignment problems depending on whose goals we consider, and analyzes the different solutions necessitated by each. The direct alignment problem considers whether an AI system accomplishes the goals of the entity operating it. In contrast, the social alignment problem considers the effects of an AI system on larger groups or on society more broadly. In particular, it also considers whether the system imposes externalities on others. Whereas solutions to the direct alignment problem center around more robust implementation, social alignment problems typically arise because of conflicts between individual and group-level goals, elevating the importance of AI governance to mediate such conflicts. Addressing the social alignment problem requires both enforcing existing norms on their developers and operators and designing new norms that apply directly to AI systems.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Hidden Agenda: a Social Deduction Game with Diverse Learned Equilibria
A key challenge in the study of multiagent cooperation is the need for individual agents not only to cooperate effectively, but to decide with whom to cooperate. This is particularly critical in situations when other age…
ToM2C: Target-oriented Multi-agent Communication and Cooperation with Theory of Mind
Being able to predict the mental states of others is a key factor to effective social interaction. It is also crucial for distributed multi-agent systems, where agents are required to communicate and cooperate. In this p…
ProToM: Promoting Prosocial Behaviour via Theory of Mind-Informed Feedback
While humans are inherently social creatures, the challenge of identifying when and how to assist and collaborate with others - particularly when pursuing independent goals - can hinder cooperation. To address this chall…
Interpretable to Whom? A Role-based Model for Analyzing Interpretable Machine Learning Systems
Several researchers have argued that a machine learning system's interpretability should be defined in relation to a specific agent or task: we should not ask if the system is interpretable, but to whom is it interpretab…
BIG-bench Machine LearningInterpretable Machine LearningRelationEmergent social conventions and collective bias in LLM populations
Social conventions are the backbone of social coordination, shaping how individuals form a group. As growing populations of artificial intelligence (AI) agents communicate through natural language, a fundamental question…
Language ModelingLanguage ModellingLarge Language Model