paper-with-me

Papers

Model evaluation for extreme risks

2023-05-24 · Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, Allan Dafoe

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is critical for addressing extreme risks. Developers must be able to identify dangerous capabilities (through "dangerous capability evaluations") and the propensity of models to apply their capabilities for harm (through "alignment evaluations"). These evaluations will become critical for keeping policymakers and other stakeholders informed, and for making responsible decisions about model training, deployment, and security.

📄 PDF Abstract BibTeX arXiv:2305.15324

Code (0)

등록된 구현이 없습니다.

Tasks

model

Similar Papers 제목 키워드 기반

Estimating value at risk and conditional tail expectation for extreme and aggregate risks

2021-01-29 · Suman Thapa, Yiqiang Q. Zhao

In this paper, we investigate risk measures such as value at risk (VaR) and the conditional tail expectation (CTE) of the extreme (maximum and minimum) and the aggregate (total) of two dependent risks. In finance, insura…

Modeling Multivariate Cyber Risks: Deep Learning Dating Extreme Value Theory

2021-03-15 · Mingyue Zhang Wu, Jinzhu Luo, Xing Fang, Maochao Xu 외

Modeling cyber risks has been an important but challenging task in the domain of cyber security. It is mainly because of the high dimensionality and heavy tails of risk patterns. Those obstacles have hindered the develop…

Deep LearningPrediction

Managing extreme AI risks amid rapid progress

2023-10-26 · Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song 외

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon …

Extreme value statistics for censored data with heavy tails under competing risks

2017-01-19 · Julien Worms, Rym Worms

This paper addresses the problem of estimating, in the presence of random censoring as well as competing risks, the extreme value index of the (sub)-distribution function associated to one particular cause, in the heavy-…

A VAE Approach to Sample Multivariate Extremes

2023-06-19 · Nicolas Lafon, Philippe Naveau, Ronan Fablet

Generating accurate extremes from an observational data set is crucial when seeking to estimate risks associated with the occurrence of future extremes which could be larger than those already observed. Applications rang…