paper-with-me

홈 › Papers

AI and Human Oversight: A Risk-Based Framework for Alignment

2025-10-10 · Laxmiraju Kandikatla, Branislav Radeljic arxiv

As Artificial Intelligence (AI) technologies continue to advance, protecting human autonomy and promoting ethical decision-making are essential to fostering trust and accountability. Human agency (the capacity of individuals to make informed decisions) should be actively preserved and reinforced by AI systems. This paper examines strategies for designing AI systems that uphold fundamental rights, strengthen human agency, and embed effective human oversight mechanisms. It discusses key oversight models, including Human-in-Command (HIC), Human-in-the-Loop (HITL), and Human-on-the-Loop (HOTL), and proposes a risk-based framework to guide the implementation of these mechanisms. By linking the level of AI model risk to the appropriate form of human oversight, the paper underscores the critical role of human involvement in the responsible deployment of AI, balancing technological innovation with the protection of individual values and rights. In doing so, it aims to ensure that AI technologies are used responsibly, safeguarding individual autonomy while maximizing societal benefits.

📄 PDF Abstract BibTeX arXiv:2510.09090

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Super Co-alignment for Sustainable Symbiotic Society

2025-04-24 · Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu 외

As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead …

Incentive-Aware AI Safety via Strategic Resource Allocation: A Stackelberg Security Games Perspective

2026-02-06 · Cheol Woo Kim, Davin Choo, Tzeh Yuan Neoh, Milind Tambe arxiv

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and …

Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning

2024-02-01 · Jitao Sang, Yuhang Wang, Jing Zhang, Yanxu Zhu 외

This paper presents a follow-up study to OpenAI's recent superalignment work on Weak-to-Strong Generalization (W2SG). Superalignment focuses on ensuring that high-level AI systems remain consistent with human values and …

Ensemble LearningIn-Context Learning

Keeping an Eye on AI: A Framework for Effective Human Oversight of AI Systems

2026-04-09 · Susanne Gaube, Markus Langer, Tim Miller, Kevin Baum 외 arxiv

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human oversight. However, notions of human ov…

The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy

2025-10-30 · William Overman, Mohsen Bayati arxiv

As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal control interface in which an agent chooses wh…