paper-with-me

홈 › Papers

Revealing the Incentive to Cause Distributional Shift

2021-09-29 · David Krueger, Tegan Maharaj, Jan Leike

Decisions made by machine learning systems have increasing influence on the world, yet it is common for machine learning algorithms to assume that no such influence exists. An example is the use of the i.i.d. assumption in content recommendation: In fact, the (choice of) content displayed can change users’ perceptions and preferences, or even drive them away, causing a shift in the distribution of users. We introduce the term auto-induced distributional shift (ADS) to describe the phenomenon of an algorithm causing change in the distribution of its own inputs. Leveraging ADS can be a means of increasing performance. But this is not always desirable, since performance metrics often underspecify what type of behaviour is desirable. When real-world conditions violate assumptions (such as i.i.d. data), this underspecification can result in unexpected behaviour. To diagnose such issues, we introduce the approach of unit tests for incentives: simple environments designed to show whether an algorithm will hide or reveal incentives to achieve performance via certain means (in our case, via ADS). We use these unit tests to demonstrate that changes to the learning algorithm (e.g. introducing meta-learning) can cause previously hidden incentives to be revealed, resulting in qualitatively different behaviour despite no change in performance metric. We further introduce a toy environment for modelling real-world issues with ADS in content recommendation, where we demonstrate that strong meta-learners achieve gains in performance via ADS. These experiments confirm that the unit tests work – an algorithm’s failure of the unit test correctly diagnoses its propensity to reveal incentives for ADS.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Meta-Learning

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Hidden incentives for self-induced distributional shift

2019-09-25 · David Scott Krueger, Tegan Maharaj, Shane Legg, Jan Leike

Decisions made by machine learning systems have increasing influence on the world. Yet it is common for machine learning algorithms to assume that no such influence exists. An example is the use of the i.i.d. assumption …

BIG-bench Machine LearningMeta-Learning

Hidden Incentives for Auto-Induced Distributional Shift

2020-09-19 · David Krueger, Tegan Maharaj, Jan Leike

Decisions made by machine learning systems have increasing influence on the world, yet it is common for machine learning algorithms to assume that no such influence exists. An example is the use of the i.i.d. assumption …

BIG-bench Machine LearningMeta-LearningQ-Learning

Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models

2025-04-07 · Jiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 외

Large language models (LLMs) are foundational explorations to artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning achieves only superficial compliance. He…

Explaining Concept Drift through the Evolution of Group Counterfactuals

2025-09-11 · Ignacy Stępka, Jerzy Stefanowski arxiv

Machine learning models in dynamic environments often suffer from concept drift, where changes in the data distribution degrade performance. While detecting this drift is a well-studied topic, explaining how and why the …

Addressing distributional shifts in operations management: The case of order fulfillment in customized production

2023-04-24 · Julian Senoner, Bernhard Kratzwald, Milan Kuzmanovic, Torbjørn H. Netland 외

To meet order fulfillment targets, manufacturers seek to optimize production schedules. Machine learning can support this objective by predicting throughput times on production lines given order specifications. However, …

Decision MakingJob Shop SchedulingManagementScheduling