paper-with-me

홈 › Papers

Monte Carlo Expected Threat (MOCET) Scoring

2025-11-20 · Joseph Kim, Saahith Potluri arxiv

Evaluating and measuring AI Safety Level (ASL) threats are crucial for guiding stakeholders to implement safeguards that keep risks within acceptable limits. ASL-3+ models present a unique risk in their ability to uplift novice non-state actors, especially in the realm of biosecurity. Existing evaluation metrics, such as LAB-Bench, BioLP-bench, and WMDP, can reliably assess model uplift and domain knowledge. However, metrics that better contextualize "real-world risks" are needed to inform the safety case for LLMs, along with scalable, open-ended metrics to keep pace with their rapid advancements. To address both gaps, we introduce MOCET, an interpretable and doubly-scalable metric (automatable and open-ended) that can quantify real-world risks.

📄 PDF Abstract BibTeX arXiv:2511.16823

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unbiased MLMC stochastic gradient-based optimization of Bayesian experimental designs

2020-05-18 · Takashi Goda, Tomohiko Hironaka, Wataru Kitade, Adam Foster

In this paper we propose an efficient stochastic optimization algorithm to search for Bayesian experimental designs such that the expected information gain is maximized. The gradient of the expected information gain with…

Experimental DesignStochastic Optimization

On-line Policy Improvement using Monte-Carlo Search

2025-01-09 · Gerald Tesauro, Gregory R. Galperin

We present a Monte-Carlo simulation algorithm for real-time policy improvement of an adaptive controller. In the Monte-Carlo simulation, the long-term expected reward of each possible action is statistically measured, us…

Training neural networks using Metropolis Monte Carlo and an adaptive variant

2022-05-16 · Stephen Whitelam, Viktor Selin, Ian Benlolo, Corneel Casert 외

We examine the zero-temperature Metropolis Monte Carlo algorithm as a tool for training a neural network by minimizing a loss function. We find that, as expected on theoretical grounds and shown empirically by other auth…

Finite Difference Solution Ansatz approach in Least-Squares Monte Carlo

2023-05-16 · Jiawei Huo

This article presents a simple but effective and efficient approach to improve the accuracy and stability of Least-Squares Monte Carlo. The key idea is to construct the ansatz of conditional expected continuation payoff …

Renewal Monte Carlo: Renewal theory based reinforcement learning

2018-04-03 · Jayakumar Subramanian, Aditya Mahajan

In this paper, we present an online reinforcement learning algorithm, called Renewal Monte Carlo (RMC), for infinite horizon Markov decision processes with a designated start state. RMC is a Monte Carlo algorithm and ret…

Managementreinforcement-learningReinforcement LearningReinforcement Learning (RL)