paper-with-me

홈 › Papers

Overseeing Agents Without Constant Oversight: Challenges and Opportunities

2026-02-18 · Madeleine Grunde-McLaughlin, Hussein Mozannar, Maya Murad, Jingya Chen, Saleema Amershi, Adam Fourney arxiv

To enable human oversight, agentic AI systems often provide a trace of reasoning and action steps. Designing traces to have an informative, but not overwhelming, level of detail remains a critical challenge. In three user studies on a Computer User Agent, we investigate the utility of basic action traces for verification, explore three alternatives via design probes, and test a novel interface's impact on error finding in question-answering tasks. As expected, we find that current practices are cumbersome, limiting their efficacy. Conversely, our proposed design reduced the time participants spent finding errors. However, although participants reported higher levels of confidence in their decisions, their final accuracy was not meaningfully improved. To this end, our study surfaces challenges for human verification of agentic systems, including managing built-in assumptions, users' subjective and changing correctness criteria, and the shortcomings, yet importance, of communicating the agent's process.

📄 PDF Abstract BibTeX arXiv:2602.16844

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Practical challenges of control monitoring in frontier AI deployments

2025-12-15 · David Lindner, Charlie Griffin, Tomek Korbak, Roland S. Zimmermann 외 arxiv

Automated control monitors could play an important role in overseeing highly capable AI agents that we do not fully trust. Prior work has explored control monitoring in simplified settings, but scaling monitoring to real…

Towards physician-centered oversight of conversational diagnostic AI

2025-07-21 · Elahe Vedadi, David Barrett, Natalie Harris, Ellery Wulczyn 외 arxiv

Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagnoses and treatment plans is considered a…

Scaling Laws For Scalable Oversight

2025-04-25 · Joshua Engels, David D. Baek, Subhash Kantamneni, Max Tegmark

Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems. However, it is still unclear how scalable oversight itse…

Chatbot

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents

2026-06-03 · Shipi Dhanorkar, Samir Passi, Mihaela Vorvoreanu arxiv

Autonomous software agents hold promise to increase developer productivity but make mistakes and exhibit novel failure modes, making human oversight central to successful human-agent collaboration. Existing research on a…

Human Oversight and Overload: Two Hidden and Costly Burdens of AI-Assisted Software Engineering

2026-06-04 · Vahid Garousi arxiv

AI is changing how software engineers work, but it often comes with hidden burdens and costs. In this paper, we characterize two such often-overlooked burdens: (1) the constant need for human oversight and inspection of …