paper-with-me

Papers

Contextual Online Pricing with (Biased) Offline Data

2025-07-03 · Yixuan Zhang, Ruihao Zhu, Qiaomin Xie arxiv

We study contextual online pricing with biased offline data. For the scalar price elasticity case, we identify the instance-dependent quantity $δ^2$ that measures how far the offline data lies from the (unknown) online optimum. We show that the time length $T$, bias bound $V$, size $N$ and dispersion $λ_{\min}(\hatΣ)$ of the offline data, and $δ^2$ jointly determine the statistical complexity. An Optimism-in-the-Face-of-Uncertainty (OFU) policy achieves a minimax-optimal, instance-dependent regret bound $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT}{λ_{\min}(\hatΣ) + (N \wedge T) δ^2})\big)$. For general price elasticity, we establish a worst-case, minimax-optimal rate $\tilde{\mathcal{O}}\big(d\sqrt{T} \wedge (V^2T + \frac{dT }{λ_{\min}(\hatΣ)})\big)$ and provide a generalized OFU algorithm that attains it. When the bias bound $V$ is unknown, we design a robust variant that always guarantees sub-linear regret and strictly improves on purely online methods whenever the exact bias is small. These results deliver the first tight regret guarantees for contextual pricing in the presence of biased offline data. Our techniques also transfer verbatim to stochastic linear bandits with biased offline data, yielding analogous bounds.

📄 PDF Abstract BibTeX arXiv:2507.02762

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms

2010-03-31 · Lihong Li, Wei Chu, John Langford, Xuanhui Wang

Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these …

News RecommendationRecommendation Systems

Online Allocation and Pricing: Constant Regret via Bellman Inequalities

2019-06-14 · Alberto Vera, Siddhartha Banerjee, Itai Gurvich

We develop a framework for designing simple and efficient policies for a family of online allocation and pricing problems, that includes online packing, budget-constrained probing, dynamic pricing, and online contextual …

Multi-Armed Bandits

Uncertainty Quantification for Demand Prediction in Contextual Dynamic Pricing

2020-03-16 · Yining Wang, Xi Chen, Xiangyu Chang, Dongdong Ge

Data-driven sequential decision has found a wide range of applications in modern operations management, such as dynamic pricing, inventory control, and assortment optimization. Most existing research on data-driven seque…

Assortment OptimizationManagementUncertainty Quantificationvalid

Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits

2026-04-27 · Zean Han, Ruihan Lin, Zezhen Ding, Jiheng Zhang arxiv

Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. Using these data can reduce early online exploration when they remain informative for the online problem. When the o…

How does online shopping affect offline price sensitivity?

2025-06-18 · Shirsho Biswas, Hema Yoganarasimhan, Haonan Zhang

The rapid rise of e-commerce has transformed consumer behavior, prompting questions about how online adoption influences offline shopping. We examine whether consumers who adopt online shopping with a retailer become mor…

counterfactualSensitivity