paper-with-me

홈 › Papers

Interactionless Inverse Reinforcement Learning: A Data-Centric Framework for Durable Alignment

2026-02-16 · Elias Malomgré, Pieter Simoens arxiv

AI alignment is growing in importance, yet many current approaches learn safety behavior by directly modifying policy parameters, entangling normative constraints with the underlying policy. This often yields opaque, difficult-to-edit alignment artifacts and reduces their reuse across models or deployments, a failure mode we term Alignment Waste. We propose Interactionless Inverse Reinforcement Learning, a framework for learning inspectable, editable, and reusable reward artifacts separately from policy optimization. We further introduce the Alignment Flywheel, a human-in-the-loop lifecycle for iteratively auditing, patching, and hardening these artifacts through automated evaluation and refinement. Together, these ideas recast alignment from a disposable training expense into a durable, verifiable engineering asset.

📄 PDF Abstract BibTeX arXiv:2602.14844

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Can Optimal Transport Improve Federated Inverse Reinforcement Learning?

2026-01-01 · David Millard, Ali Baheri arxiv

In robotics and multi-agent systems, fleets of autonomous agents often operate in subtly different environments while pursuing a common high-level objective. Directly pooling their data to learn a shared reward function …

Reinforcement LearningFederated Learning

Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs

2024-04-22 · Lili Wu, Ben Evans, Riashat Islam, Raihan Seraj 외

Discovering an informative, or agent-centric, state representation that encodes only the relevant information while discarding the irrelevant is a key challenge towards scaling reinforcement learning algorithms and effic…

Representation Learning

Learning Navigation Subroutines from Egocentric Videos

2019-05-29 · Ashish Kumar, Saurabh Gupta, Jitendra Malik

Planning at a higher level of abstraction instead of low level torques improves the sample efficiency in reinforcement learning, and computational efficiency in classical planning. We propose a method to learn such hiera…

Computational EfficiencyPseudo LabelReinforcement Learning

GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement Learning

2025-08-14 · Kelin Yu, Sheng Zhang, Harshit Soora, Furong Huang 외 arxiv

Recent advances have shown that video generation models can enhance robot learning by deriving effective robot actions through inverse dynamics. However, these methods heavily depend on the quality of generated data and …

Reinforcement LearningVideo Generation

Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation

2025-12-05 · Fabian Konstantinidis, Moritz Sackmann, Ulrich Hofmann, Christoph Stiller arxiv

Scalable multi-agent driving simulation requires behavior models that are both realistic and computationally efficient. We address this by optimizing the behavior model that controls individual traffic participants. To i…

Reinforcement Learning