paper-with-me

Papers

Quantifying Misalignment Between Agents: Towards a Sociotechnical Understanding of Alignment

2024-06-06 · Aidan Kierans, Avijit Ghosh, Hananel Hazan, Shiri Dori-Hacohen

Existing work on the alignment problem has focused mainly on (1) qualitative descriptions of the alignment problem; (2) attempting to align AI actions with human interests by focusing on value specification and learning; and/or (3) focusing on a single agent or on humanity as a monolith. Recent sociotechnical approaches highlight the need to understand complex misalignment among multiple human and AI agents. We address this gap by adapting a computational social science model of human contention to the alignment problem. Our model quantifies misalignment in large, diverse agent groups with potentially conflicting goals across various problem areas. Misalignment scores in our framework depend on the observed agent population, the domain in question, and conflict between agents' weighted preferences. Through simulations, we demonstrate how our model captures intuitive aspects of misalignment across different scenarios. We then apply our model to two case studies, including an autonomous vehicle setting, showcasing its practical utility. Our approach offers enhanced explanatory power for complex sociotechnical environments and could inform the design of more aligned AI systems in real-world applications.

📄 PDF Abstract BibTeX arXiv:2406.04231

Code (1)

riet-lab/quantifying-misalignment 공식 구현

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Roadmap on Incentive Compatibility for AI Alignment and Governance in Sociotechnical Systems

2024-02-20 · Zhaowei Zhang, Fengshuo Bai, Mingzhi Wang, Haoyang Ye 외

The burgeoning integration of artificial intelligence (AI) into human society brings forth significant implications for societal governance and safety. While considerable strides have been made in addressing AI alignment…

Charting the Sociotechnical Gap in Explainable AI: A Framework to Address the Gap in XAI

2023-02-01 · Upol Ehsan, Koustuv Saha, Munmun De Choudhury, Mark O. Riedl

Explainable AI (XAI) systems are sociotechnical in nature; thus, they are subject to the sociotechnical gap--divide between the technical affordances and the social needs. However, charting this gap is challenging. In th…

Explainable Artificial Intelligence (XAI)

Dimensions of Disagreement: Unpacking Divergence and Misalignment in Cognitive Science and Artificial Intelligence

2023-10-03 · Kerem Oktar, Ilia Sucholutsky, Tania Lombrozo, Thomas L. Griffiths

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this lar…

Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents

2025-09-29 · Boxuan Zhang, Yi Yu, Jiaxuan Guo, Jing Shao arxiv

The prevalent deployment of Large Language Model agents such as OpenClaw unlocks potential in real-world applications, while amplifying safety concerns. Among these concerns, the self-replication risk of LLM agents drive…

A Field Guide to Deploying AI Agents in Clinical Practice

2025-09-30 · Jack Gallifant, Katherine C. Kellogg, Matt Butler, Amanda Centi 외 arxiv

Large language models (LLMs) integrated into agent-driven workflows hold immense promise for healthcare, yet a significant gap exists between their potential and practical implementation within clinical settings. To addr…

Prompt Engineering