paper-with-me

홈 › Papers

Progressive Autonomy as Preference Learning: A Formalization of Trust Calibration for Agentic Tool Use

2026-05-18 · Changkun Ou arxiv

We formalize trust calibration for agentic tool use (deciding when an automated agent's proposed action may execute autonomously versus require human approval) as a preference-learning problem. A policy gateway maintains a Gaussian-process posterior over a latent human risk-tolerance function, observed through a probit likelihood on binary approve/deny feedback, and escalates to the human exactly where the approval outcome is most uncertain. We show this is structurally an instance of Preferential Bayesian Optimization, inheriting its inference machinery (approximate Gaussian-process classification) and its sample-efficiency argument (uncertainty-targeted querying), while differing in objective: classifying an action space into allow/block/ask regions rather than optimizing a design.

📄 PDF Abstract BibTeX arXiv:2605.19151

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomy

2025-05-29 · Ahmad Mohsin, Helge Janicke, Ahmed Ibrahim, Iqbal H. Sarker 외

This article presents a structured framework for Human-AI collaboration in Security Operations Centers (SOCs), integrating AI autonomy, trust calibration, and Human-in-the-loop decision making. Existing frameworks in SOC…

Decision Making

What Is My Robot Thinking? Design Considerations for Transparent and Trustworthy Shared Autonomy

2026-06-05 · Atharv Belsare, Zohre Karimi, Connor Mattson, Rushiil Nakka 외 arxiv

Assistive robots operating under shared autonomy must balance user control with autonomous assistance. Because robot actions depend on internal intent inference that is not directly observable, mismatches between inferre…

Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI

2024-12-08 · Chao Yang, Chaochao Lu, Yingchun Wang, BoWen Zhou

Ensuring Artificial General Intelligence (AGI) reliably avoids harmful behaviors is a critical challenge, especially for systems with high autonomy or in safety-critical domains. Despite various safety assurance proposal…

Decision Making

From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy

2026-04-29 · Serhii Zabolotnii, Viktoriia Holinko, Olha Antonenko arxiv

Trust in clinical artificial intelligence (AI) cannot be reduced to model accuracy, fluency of generation, or overall positive user impression. In medicine, trust must be engineered as a measurable system property ground…

Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration

2026-04-07 · Yi Yuan, Xuhong Wang, Shanzhe Lei arxiv

As agent-based systems continue to evolve, deep research agents are capable of automatically generating research-style reports across diverse domains. While these agents promise to streamline information synthesis and kn…