paper-with-me

홈 › Papers

APC-RL: Exceeding Data-Driven Behavior Priors with Adaptive Policy Composition

2026-01-27 · Finn Rietz, Pedro Zuidberg dos Martires, Johannes Andreas Stork arxiv

Incorporating demonstration data into reinforcement learning (RL) can greatly accelerate learning, but existing approaches often assume demonstrations are optimal and fully aligned with the target task. In practice, demonstrations are frequently sparse, suboptimal, or misaligned, which can degrade performance when these demonstrations are integrated into RL. We propose Adaptive Policy Composition (APC), a hierarchical model that adaptively composes multiple data-driven Normalizing Flow (NF) priors. Instead of enforcing strict adherence to the priors, APC estimates each prior's applicability to the target task while leveraging them for exploration. Moreover, APC either refines useful priors, or sidesteps misaligned ones when necessary to optimize downstream reward. Across diverse benchmarks, APC accelerates learning when demonstrations are aligned, remains robust under severe misalignment, and leverages suboptimal demonstrations to bootstrap exploration while avoiding performance degradation caused by overly strict adherence to suboptimal demonstrations.

📄 PDF Abstract BibTeX arXiv:2601.19452

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces

2026-02-02 · Mona Rajhans arxiv

Modern cybersecurity platforms must process and display high-frequency telemetry such as network logs, endpoint events, alerts, and policy changes in real time. Traditional rendering techniques based on static pagination…

HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion

2026-05-08 · Vickson Ferrel arxiv

Insider threat detection assumes that an adaptive insider leaves behavioral residue distinguishing them from legitimate users. We test this assumption against an LLM-driven adaptive insider in a controlled multi-agent si…

DSPR: Dual-Stream Physics-Residual Networks for Trustworthy Industrial Time Series Forecasting

2026-04-08 · Yeran Zhang, Pengwei Yang, Guoqing Wang, Tianyu Li arxiv

Accurate forecasting of industrial time series requires balancing predictive accuracy with physical plausibility under non-stationary operating conditions. Existing data-driven models often achieve strong statistical per…

Time Series Forecasting

Active Inference with Reusable State-Dependent Value Profiles

2025-12-03 · Jacob Poschl arxiv

Adaptive behavior in volatile environments requires agents to switch among value-control regimes across latent contexts, but maintaining separate preferences, policy biases, and action-confidence parameters for every sit…

AGMA: Adaptive Gaussian Mixture Anchors for Prior-Guided Multimodal Human Trajectory Forecasting

2026-02-04 · Chao Li, Rui Zhang, Siyuan Huang, Xian Zhong 외 arxiv

Human trajectory forecasting requires capturing the multimodal nature of pedestrian behavior. However, existing approaches suffer from prior misalignment. Their learned or fixed priors often fail to capture the full dist…

Trajectory Forecasting