paper-with-me

Papers

Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling

2021-06-19 · Mengdi Xu, Peide Huang, Fengpei Li, Jiacheng Zhu, Xuewei Qi, Kentaro Oguchi, Zhiyuan Huang, Henry Lam, Ding Zhao

Evaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibitively large number of testing iterations. On the other hand, a biased or inaccurate policy evaluation in a safety-critical system could potentially cause unexpected catastrophic failures during deployment. This paper proposes the Accelerated Policy Evaluation (APE) method, which simultaneously uncovers rare events and estimates the rare event probability in Markov decision processes. The APE method treats the environment nature as an adversarial agent and learns towards, through adaptive importance sampling, the zero-variance sampling distribution for the policy evaluation. Moreover, APE is scalable to large discrete or continuous spaces by incorporating function approximators. We investigate the convergence property of APE in the tabular setting. Our empirical studies show that APE can estimate the rare event probability with a smaller bias while only using orders of magnitude fewer samples than baselines in multi-agent and single-agent environments.

📄 PDF Abstract BibTeX arXiv:2106.10566

Code (1)

eleurent/highway-env 공식 구현

Similar Papers 제목 키워드 기반

Deep Probabilistic Accelerated Evaluation: A Robust Certifiable Rare-Event Simulation Methodology for Black-Box Safety-Critical Systems

2020-06-28 · Mansur Arief, Zhiyuan Huang, Guru Koushik Senthil Kumar, Yuanlu Bai 외

Evaluating the reliability of intelligent physical systems against rare safety-critical events poses a huge testing burden for real-world applications. Simulation provides a useful platform to evaluate the extremal risks…

Policy-Grounded Safety Evaluation of 20 Large Language Models

2025-07-19 · Juan Manuel Contreras arxiv

As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI, a programmatic platform for generating a…

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

2026-02-12 · Keita Broadwater arxiv

Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation across diverse tasks and risk categories. However, real-world deployment of…

Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents

2025-09-28 · Gauri Kholkar, Ratinder Ahuja arxiv

As autonomous AI agents are used in regulated and safety-critical settings, organizations need effective ways to turn policy into enforceable controls. We introduce a regulatory machine learning framework that converts u…

SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation

2025-10-24 · Alec Helbling, Shruti Palaskar, Kundan Krishna, Polo Chau 외 arxiv

What exactly makes a particular image unsafe? Systematically differentiating between benign and problematic images is a challenging problem, as subtle changes to an image, such as an insulting gesture or symbol, can dras…

Data AugmentationImage GenerationImage Editing