paper-with-me

홈 › Papers

Diffuse AI Control on Fuzzy Tasks

2026-06-08 · Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton arxiv

AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned with mitigating risks from AI sabotage distributed over long deployment horizons (diffuse threats). These risks are particularly pernicious on fuzzy tasks, i.e. tasks which are hard to grade or require intuition. To understand diffuse threats on fuzzy tasks, we introduce a framework that considers AI control as an adversarial game between a blue team and a red team. The blue team uses a weak trusted model to construct a weak score against which they would train a strong, potentially subversive model to remove the subversion propensity if it were present. The red team then tries to find model behaviors that are rated highly by the weak score, and thus might not be trained out, but actually correspond to poor performance. We test our framework on the task of writing experimental proposals for research questions from recent ML papers. We use a language model with access to the original paper as a proxy "ground-truth" scorer. Our red team discovers subversive behaviors using multi-objective evolutionary prompt optimization. We show that Opus~4.6 can write proposals that are worse according to the ground truth proxy than those of GPT-OSS-20B, while the weak scorer rates them as highly as the best proposals from Opus 4.6. We then propose an adversarial optimization algorithm for the blue team that discovers more robust prompts for the weak model. This algorithm produces a blue team prompt that our red team optimization fails to exploit.

📄 PDF Abstract BibTeX arXiv:2606.08892

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Neuro-Fuzzy Inference System and a Multilayer Perceptron Model Trained with Grey Wolf Optimizer for Predicting Solar Diffuse Fraction

2020-09-13 · Randall Claywell, Laszlo Nadai, Felde Imre, Amir Mosavi

The accurate prediction of the solar Diffuse Fraction (DF), sometimes called the Diffuse Ratio, is an important topic for solar energy research. In the present study, the current state of Diffuse Irradiance research is d…

Distilling Deep RL Models Into Interpretable Neuro-Fuzzy Systems

2022-09-07 · Arne Gevaert, Jonathan Peck, Yvan Saeys

Deep Reinforcement Learning uses a deep neural network to encode a policy, which achieves very good performance in a wide range of applications but is widely regarded as a black box model. A more interpretable alternativ…

Deep Reinforcement LearningOpenAI Gymreinforcement-learningReinforcement Learning+1

A Hierarchical Genetic Optimization of a Fuzzy Logic System for Flow Control in Micro Grids

2016-04-16 · Enrico De Santis, Antonello Rizzi, Alireza Sadeghian

Bio-inspired algorithms like Genetic Algorithms and Fuzzy Inference Systems (FIS) are nowadays widely adopted as hybrid techniques in commercial and industrial environment. In this paper we present an interesting applica…

Decision Makingenergy tradingManagement

Constrained Diffusers for Safe Planning and Control

2025-06-14 · Jichen Zhang, Liqun Zhao, Antonis Papachristodoulou, Jack Umenberger

Diffusion models have shown remarkable potential in planning and control tasks due to their ability to represent multimodal distributions over actions and trajectories. However, ensuring safety under constraints remains …

Optimization of Fuzzy Controller of a Wind Power Plant Based on the Swarm Intelligence

2020-06-14 · Vadim Manusov, Pavel Matrenin

The article considers the problem of the optimal control of a wind power plant based on fuzzy control and automation of generating the fuzzy rule base. Fuzzy rules by experts do not always provide a maximum power output …