paper-with-me

Papers

Off-Policy Evaluation for Sequential Persuasion Process with Unobserved Confounding

2025-04-01 · Nishanth Venkatesh S., Heeseung Bang, Andreas A. Malikopoulos

In this paper, we expand the Bayesian persuasion framework to account for unobserved confounding variables in sender-receiver interactions. While traditional models assume that belief updates follow Bayesian principles, real-world scenarios often involve hidden variables that impact the receiver's belief formation and decision-making. We conceptualize this as a sequential decision-making problem, where the sender and receiver interact over multiple rounds. In each round, the sender communicates with the receiver, who also interacts with the environment. Crucially, the receiver's belief update is affected by an unobserved confounding variable. By reformulating this scenario as a Partially Observable Markov Decision Process (POMDP), we capture the sender's incomplete information regarding both the dynamics of the receiver's beliefs and the unobserved confounder. We prove that finding an optimal observation-based policy in this POMDP is equivalent to solving for an optimal signaling strategy in the original persuasion framework. Furthermore, we demonstrate how this reformulation facilitates the application of proximal learning for off-policy evaluation in the persuasion process. This advancement enables the sender to evaluate alternative signaling strategies using only observational data from a behavioral policy, thus eliminating the necessity for costly new experiments.

📄 PDF Abstract BibTeX arXiv:2504.01211

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingOff-policy evaluationSequential Decision Making

Similar Papers 제목 키워드 기반

Off-policy Policy Evaluation For Sequential Decisions Under Unobserved Confounding

2020-03-12 · NeurIPS 2020 12 · Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma Brunskill

When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluation policies before deploying them. This…

Decision MakingManagementSequential Decision Making

Robust Fitted-Q-Evaluation and Iteration under Sequentially Exogenous Unobserved Confounders

2023-02-01 · David Bruns-Smith, Angela Zhou

Offline reinforcement learning is important in domains such as medicine, economics, and e-commerce where online experimentation is costly, dangerous or unethical, and where the true model is unknown. However, most method…

reinforcement-learningReinforcement LearningSensitivityvalid

Markov Persuasion Processes: Learning to Persuade from Scratch

2024-02-05 · Francesco Bacchiocchi, Francesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi 외

In Bayesian persuasion, an informed sender strategically discloses information to a receiver so as to persuade them to undertake desirable actions. Recently, a growing attention has been devoted to settings in which send…

Persuasiveness

Sequential Information Design: Markov Persuasion Process and Its Efficient Reinforcement Learning

2022-02-22 · Jibang Wu, Zixuan Zhang, Zhe Feng, Zhaoran Wang 외

In today's economy, it becomes important for Internet platforms to consider the sequential information design problem to align its long term interest with incentives of the gig service providers. This paper proposes a no…

reinforcement-learningReinforcement Learning (RL)

Confounding-Robust Policy Evaluation in Infinite-Horizon Reinforcement Learning

2020-02-11 · NeurIPS 2020 12 · Nathan Kallus, Angela Zhou

Off-policy evaluation of sequential decision policies from observational data is necessary in applications of batch reinforcement learning such as education and healthcare. In such settings, however, unobserved variables…

Off-policy evaluationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1