paper-with-me

홈 › Papers

Incorporating Behavioral Constraints in Online AI Systems

2018-09-15 · Avinash Balakrishnan, Djallel Bouneffouf, Nicholas Mattei, Francesca Rossi

AI systems that learn through reward feedback about the actions they take are increasingly deployed in domains that have significant impact on our daily life. However, in many cases the online rewards should not be the only guiding criteria, as there are additional constraints and/or priorities imposed by regulations, values, preferences, or ethical principles. We detail a novel online agent that learns a set of behavioral constraints by observation and uses these learned constraints as a guide when making decisions in an online setting while still being reactive to reward feedback. To define this agent, we propose to adopt a novel extension to the classical contextual multi-armed bandit setting and we provide a new algorithm called Behavior Constrained Thompson Sampling (BCTS) that allows for online learning while obeying exogenous constraints. Our agent learns a constrained policy that implements the observed behavioral constraints demonstrated by a teacher agent, and then uses this constrained policy to guide the reward-based online exploration and exploitation. We characterize the upper bound on the expected regret of the contextual bandit algorithm that underlies our agent and provide a case study with real world data in two application domains. Our experiments show that the designed agent is able to act within the set of behavior constraints without significantly degrading its overall reward performance.

📄 PDF Abstract BibTeX arXiv:1809.05720

Code (0)

등록된 구현이 없습니다.

Tasks

Thompson Sampling

Similar Papers 제목 키워드 기반

BRAID: Input-Driven Nonlinear Dynamical Modeling of Neural-Behavioral Data

2025-09-23 · Parsa Vahidi, Omid G. Sani, Maryam M. Shanechi arxiv

Neural populations exhibit complex recurrent structures that drive behavior, while continuously receiving and integrating external inputs from sensory stimuli, upstream regions, and neurostimulation. However, neural popu…

A Behavioral Framework for Data-Driven Modeling of Nonlinear Systems in Vector-Valued Reproducing Kernel Hilbert Spaces

2026-05-08 · Boya Hou, Maxim Raginsky arxiv

We generalize Jan Willems' behavioral approach to a class of discrete-time nonlinear systems in a vector-valued reproducing kernel Hilbert space (RKHS). Apart from linear time-invariant systems, this class covers nonline…

Are they human? Detecting large language models by probing human memory constraints

2026-03-10 · Simon Schug, Brenden M. Lake arxiv

The validity of online behavioral research relies on study participants being human rather than machine. In the past, it was possible to detect machines by posing simple challenges that were easily solved by humans but n…

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems

2026-05-15 · Parand A. Alamdari, Toryn Q. Klassen, Sheila A. McIlraith arxiv

We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle, from pre-deployment testing to post-deployment auditing. Combining …

Data-Driven Predictive Control for Linear Parameter-Varying Systems

2021-03-30 · Chris Verhoek, Hossam S. Abbas, Roland Tóth, Sofie Haesaert

Based on the extension of the behavioral theory and the Fundamental Lemma for Linear Parameter-Varying (LPV) systems, this paper introduces a Data-driven Predictive Control (DPC) scheme capable to ensure reference tracki…

LEMMAScheduling