paper-with-me

Papers

Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors

2025-09-15 · Anirudha Majumdar arxiv

This paper proposes deception as a mechanism for out-of-distribution (OOD) generalization: by learning data representations that make training data appear independent and identically distributed (iid) to an observer, we can identify stable features that eliminate spurious correlations and generalize to unseen domains. We refer to this principle as deceptive risk minimization (DRM) and instantiate it with a practical differentiable objective that simultaneously learns features that eliminate distribution shifts from the perspective of a detector based on conformal martingales while minimizing a task-specific loss. In contrast to domain adaptation or prior invariant representation learning methods, DRM does not require access to test data or a partitioning of training data into a finite number of data-generating domains. We demonstrate the efficacy of DRM on numerical experiments with concept shift and a simulated imitation learning setting with covariate shift in environments that a robot is deployed in.

📄 PDF Abstract BibTeX arXiv:2509.12081

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningDomain Adaptation

Similar Papers 제목 키워드 기반

Deceptive Decision-Making Under Uncertainty

2021-09-14 · Yagiz Savas, Christos K. Verginis, Ufuk Topcu

We study the design of autonomous agents that are capable of deceiving outside observers about their intentions while carrying out tasks in stochastic, complex environments. By modeling the agent's behavior as a Markov d…

Decision MakingDecision Making Under Uncertainty

Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant

2024-04-25 · Olli Järviniemi, Evan Hubinger

We study the tendency of AI systems to deceive by constructing a realistic simulation setting of a company AI assistant. The simulated company employees provide tasks for the assistant to complete, these tasks spanning w…

Information Retrieval

Invariant Risk Minimization

2019-07-05 · Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, David Lopez-Paz

We introduce Invariant Risk Minimization (IRM), a learning paradigm to estimate invariant correlations across multiple training distributions. To achieve this goal, IRM learns a data representation such that the optimal …

Domain GeneralizationImage ClassificationOut-of-Distribution Generalization

Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems

2025-04-10 · Simon Lermen, Mateusz Dziemian, Natalia Pérez-Campanero Antolín

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental framework, we show that language models (…

Frustratingly Easy Model Generalization by Dummy Risk Minimization

2023-08-04 · Juncheng Wang, Jindong Wang, Xixu Hu, Shujun Wang 외

Empirical risk minimization (ERM) is a fundamental machine learning paradigm. However, its generalization ability is limited in various tasks. In this paper, we devise Dummy Risk Minimization (DuRM), a frustratingly easy…

modelOut-of-Distribution GeneralizationSemantic Segmentation