paper-with-me

홈 › Papers

The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation

2025-05-29 · Adrien Majka, El-Mahdi El-Mhamdi

Goodhart's law is a famous adage in policy-making that states that ``When a measure becomes a target, it ceases to be a good measure''. As machine learning models and the optimisation capacity to train them grow, growing empirical evidence reinforced the belief in the validity of this law without however being formalised. Recently, a few attempts were made to formalise Goodhart's law, either by categorising variants of it, or by looking at how optimising a proxy metric affects the optimisation of an intended goal. In this work, we alleviate the simplifying independence assumption, made in previous works, and the assumption on the learning paradigm made in most of them, to study the effect of the coupling between the proxy metric and the intended goal on Goodhart's law. Our results show that in the case of light tailed goal and light tailed discrepancy, dependence does not change the nature of Goodhart's effect. However, in the light tailed goal and heavy tailed discrepancy case, we exhibit an example where over-optimisation occurs at a rate inversely proportional to the heavy tailedness of the discrepancy between the goal and the metric. %

📄 PDF Abstract BibTeX arXiv:2505.23445

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Goodhart's law, with an application to value alignment

2024-10-12 · El-Mahdi El-Mhamdi, Lê-Nguyên Hoang

``When a measure becomes a target, it ceases to be a good measure'', this adage is known as {\it Goodhart's law}. In this paper, we investigate formally this law and prove that it critically depends on the tail distribut…

Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization

2025-10-03 · Antoine Maier, Aude Maier, Tom David arxiv

A common but rarely examined assumption in machine learning is that training yields models that actually satisfy their specified objective function. We call this the Objective Satisfaction Assumption (OSA). Although devi…

Benign Oscillation of Stochastic Gradient Descent with Large Learning Rates

2023-10-26 · Miao Lu, Beining Wu, Xiaodong Yang, Difan Zou

In this work, we theoretically investigate the generalization properties of neural networks (NN) trained by stochastic gradient descent (SGD) algorithm with large learning rates. Under such a training regime, our finding…

Provable Weak-to-Strong Generalization via Benign Overfitting

2024-10-06 · David X. Wu, Anant Sahai

The classic teacher-student model in machine learning posits that a strong teacher supervises a weak student to improve the student's capabilities. We instead consider the inverted situation, where a weak teacher supervi…

Categorizing Variants of Goodhart's Law

2018-03-13 · David Manheim, Scott Garrabrant

There are several distinct failure modes for overoptimization of systems on the basis of metrics. This occurs when a metric which can be used to improve a system is used to an extent that further optimization is ineffect…