paper-with-me

홈 › Papers

Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents

2026-06-11 · Laxmipriya Ganesh Iyer, Rahul Suresh Babu arxiv

Modern decision systems increasingly rely on learned components whose outputs may be confident yet wrong, exposing downstream actions to costly errors. We introduce Risk-Aware Causal Gating (RACG), a framework that decides whether to act on, defer, or abstain from a model's prediction by combining causal effect estimation with calibrated risk control. RACG models the causal pathway from candidate actions to outcomes and gates each decision according to an estimated counterfactual risk rather than raw predictive confidence. To make gating reliable, we derive distribution-free bounds on the probability of acting under high-risk conditions and show how these bounds translate into operating thresholds that satisfy user-specified safety constraints. We further propose an adaptive gating policy that adjusts to distribution shift by monitoring discrepancies between predicted and realized outcomes, tightening the gate when causal assumptions appear violated. Across simulated interventions and real-world decision benchmarks, RACG reduces high-cost errors substantially while preserving most of the utility of an ungated policy, and it outperforms confidence-based and selective-prediction baselines at matched abstention rates. Our results indicate that explicitly separating causal risk from predictive uncertainty yields decision systems that are both safer and more transparent, offering a principled mechanism for trustworthy automation in high-stakes settings.

📄 PDF Abstract BibTeX arXiv:2606.13884

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

2026-07-29 · Marc Kaufeld, Dian Zhuang, Johannes Betz arxiv

This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The proposed approach combines RBFN-based candidate trajectory generation wit…

Autonomous DrivingMotion Planning

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

2026-06-15 · Hankyul Baek, Jaewon Noh, Sang Seo, Yongsu Kim 외 arxiv

AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools where they can read, update, and disseminate sensitive information. Much of prior …

Harmonizing AI Safety Thresholds

2026-07-17 · Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza, Markov Grey arxiv

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. More…

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety

2026-06-26 · Ting Ma, Xiufeng Huang, Benlei Cui, Xiaowen Xu 외 arxiv

As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safety is adversarial: many failures arise no…

Adversarial RobustnessReinforcement Learning

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

2024-01-18 · Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang 외

Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive …

Benchmarking