paper-with-me

홈 › Papers

Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures

2018-12-04 · ICLR 2019 5 · Jonathan Uesato, Ananya Kumar, Csaba Szepesvari, Tom Erez, Avraham Ruderman, Keith Anderson, Krishmamurthy, Dvijotham, Nicolas Heess, Pushmeet Kohli

This paper addresses the problem of evaluating learning systems in safety critical domains such as autonomous driving, where failures can have catastrophic consequences. We focus on two problems: searching for scenarios when learned agents fail and assessing their probability of failure. The standard method for agent evaluation in reinforcement learning, Vanilla Monte Carlo, can miss failures entirely, leading to the deployment of unsafe agents. We demonstrate this is an issue for current agents, where even matching the compute used for training is sometimes insufficient for evaluation. To address this shortcoming, we draw upon the rare event probability estimation literature and propose an adversarial evaluation approach. Our approach focuses evaluation on adversarially chosen situations, while still providing unbiased estimates of failure probabilities. The key difficulty is in identifying these adversarial situations -- since failures are rare there is little signal to drive optimization. To solve this we propose a continuation approach that learns failure modes in related but less robust agents. Our approach also allows reuse of data already collected for training the agent. We demonstrate the efficacy of adversarial evaluation on two standard domains: humanoid control and simulated driving. Experimental results show that our methods can find catastrophic failures and estimate failures rates of agents multiple orders of magnitude faster than standard evaluation schemes, in minutes to hours rather than days.

📄 PDF Abstract BibTeX arXiv:1812.01647

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingHumanoid ControlReinforcement Learning

Similar Papers 제목 키워드 기반

Scalable Safety-Critical Policy Evaluation with Accelerated Rare Event Sampling

2021-06-19 · Mengdi Xu, Peide Huang, Fengpei Li, Jiacheng Zhu 외

Evaluating rare but high-stakes events is one of the main challenges in obtaining reliable reinforcement learning policies, especially in large or infinite state/action spaces where limited scalability dictates a prohibi…

Realistic Adversarial Attacks for Robustness Evaluation of Trajectory Prediction Models via Future State Perturbation

2025-05-09 · Julian F. Schumann, Jeroen Hagenus, Frederik Baymler Mathiesen, Arkady Zgonnikov

Trajectory prediction is a key element of autonomous vehicle systems, enabling them to anticipate and react to the movements of other road users. Evaluating the robustness of prediction models against adversarial attacks…

Autonomous VehiclesPredictionTrajectory Prediction

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

2025-11-29 · Jiacheng Guo, Suozhi Huang, Zixin Yao, Yifan Zhang 외 arxiv

This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in the uniquely demanding and fast-paced cry…

DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery

2025-08-09 · Keyu Li, Mohan Jiang, Dayuan Fu, Yunze Wu 외 arxiv

The rapid advancement of large language models has fundamentally shifted the bottleneck in AI development from computational power to data availability-with countless valuable datasets remaining hidden across specialized…

Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities

2024-10-10 · Andrey Anurin, Jonathan Ng, Kibo Schaffer, Jason Schreiber 외

LLM agents have the potential to revolutionize defensive cyber operations, but their offensive capabilities are not yet fully understood. To prepare for emerging threats, model developers and governments are evaluating t…