Contextual Online Pricing with (Biased) Offline Data
We study contextual online pricing with biased offline data. For the scalar price elasticity case, we identify the instance-dependent quantity $δ^2$ that measures how far the offline data lies from the (unknown) online optimum. We show that the time length $T$, bias bound $V$, size $N$ and dispersion $λ_{\min}(\hatΣ)$ of the offline data, and $δ^2$ jointly determine the statistical complexity. An Optimism-in-the-Face-of-Uncertainty (OFU) policy achieves a minimax-optimal, instance-dependent regret bound $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT}{λ_{\min}(\hatΣ) + (N \wedge T) δ^2})\big)$. For general price elasticity, we establish a worst-case, minimax-optimal rate $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT }{λ_{\min}(\hatΣ)})\big)$ and provide a generalized OFU algorithm that attains it. When the bias bound $V$ is unknown, we design a robust variant that always guarantees sub-linear regret and strictly improves on purely online methods whenever the exact bias is small. These results deliver the first tight regret guarantees for contextual pricing in the presence of biased offline data. Our techniques also transfer verbatim to stochastic linear bandits with biased offline data, yielding analogous bounds.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these …
News RecommendationRecommendation SystemsOnline Allocation and Pricing: Constant Regret via Bellman Inequalities
We develop a framework for designing simple and efficient policies for a family of online allocation and pricing problems, that includes online packing, budget-constrained probing, dynamic pricing, and online contextual …
Multi-Armed BanditsUncertainty Quantification for Demand Prediction in Contextual Dynamic Pricing
Data-driven sequential decision has found a wide range of applications in modern operations management, such as dynamic pricing, inventory control, and assortment optimization. Most existing research on data-driven seque…
Assortment OptimizationManagementUncertainty QuantificationvalidDirection-Aware Offline-to-Online Learning in Linear Contextual Bandits
Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. Using these data can reduce early online exploration when they remain informative for the online problem. When the o…
How does online shopping affect offline price sensitivity?
The rapid rise of e-commerce has transformed consumer behavior, prompting questions about how online adoption influences offline shopping. We examine whether consumers who adopt online shopping with a retailer become mor…
counterfactualSensitivity