paper-with-me

Papers

Learning Personalized Ad Impact via Contextual Reinforcement Learning under Delayed Rewards

2025-10-22 · Yuwei Cheng, Zifeng Zhao, Haifeng Xu arxiv

Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed and long-term effects, cumulative ad impacts such as reinforcement or fatigue, and customer heterogeneity. However, these effects are often not jointly addressed in previous studies. To capture these factors, we model ad bidding as a Contextual Markov Decision Process (CMDP) with delayed Poisson rewards. For efficient estimation, we propose a two-stage maximum likelihood estimator combined with data-splitting strategies, ensuring controlled estimation error based on the first-stage estimator's (in)accuracy. Building on this, we design a reinforcement learning algorithm to derive efficient personalized bidding strategies. This approach achieves a near-optimal regret bound of $\tilde{O}{(dH^2\sqrt{T})}$, where $d$ is the contextual dimension, $H$ is the number of rounds, and $T$ is the number of customers. Our theoretical findings are validated by simulation experiments.

📄 PDF Abstract BibTeX arXiv:2510.20055

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Bi-Level Contextual Bandits for Individualized Resource Allocation under Delayed Feedback

2025-11-13 · Mohammadsina Almasi, Hadis Anahideh arxiv

Equitably allocating limited resources in high-stakes domains-such as education, employment, and healthcare-requires balancing short-term utility with long-term impact, while accounting for delayed outcomes, hidden heter…

Reinforcement Learning from Delayed Observations via World Models

2024-03-18 · Armin Karamzade, KyungMin Kim, Montek Kalsi, Roy Fox

In standard reinforcement learning settings, agents typically assume immediate feedback about the effects of their actions after taking them. However, in practice, this assumption may not hold true due to physical constr…

continuous-controlContinuous Controlreinforcement-learningReinforcement Learning

Optimism and Delays in Episodic Reinforcement Learning

2021-11-15 · Benjamin Howson, Ciara Pike-Burke, Sarah Filippi

There are many algorithms for regret minimisation in episodic reinforcement learning. This problem is well-understood from a theoretical perspective, providing that the sequences of states, actions and rewards associated…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Cross-Contextual Vision-Language Adaptation with LoRA for Personalized Severe Adverse Event Detection in Clinical Wound Monitoring

2026-07-06 · Aditi Naiknaware, Jian Sun, Aminreza Khandan, Shengyang Huang 외 arxiv

Wound monitoring is a critical yet underserved clinical challenge, where timely identification of severe adverse events (SAEs) such as infection, tissue deterioration, and delayed healing can significantly impact patient…

Multimodal Reasoning

Budgeted Recommendation with Delayed Feedback

2024-05-19 · Kweiguu Liu, Setareh Maghsudi

In a conventional contextual multi-armed bandit problem, the feedback (or reward) is immediately observable after an action. Nevertheless, delayed feedback arises in numerous real-life situations and is particularly cruc…

Decision MakingMulti-Armed Bandits