On The Limited Cost of Consistency In Points-Based Rewards Programs
Points-based rewards programs are a prevalent way to incentivize customer loyalty. In theory, the effective design of these programs is shaped by two important considerations. Customer heterogeneity suggests that tailoring thresholds to different customer types may yield revenue gains, while demand uncertainty creates incentives to update thresholds over time as the seller learns how customers respond to the program. In practice, however, there may be nontrivial operational and customer-facing benefits associated with maintaining consistency along both of these dimensions. Setting a single threshold across customer types can facilitate transparent communication and reduce the legal, reputational, and operational risks associated with personalization, while avoiding threshold increases prevents the devaluation of customers' previously earned points. Motivated by these considerations, this paper studies the cost of consistency in points-based rewards programs. We study this question within the context of the popular ``Buy $N$, Get One Free'' (BNGO) rewards programs. We first bound the value of personalization of these programs, proving that the revenue of the optimal single-threshold program is within a factor of $\frac{π^2}{6}$ of that of the optimal personalized strategy. We then tackle the problem of designing devaluation-free learning algorithms --- i.e., algorithms that never increase the redemption threshold --- in the presence of demand uncertainty. Toward this goal, we first design a greedy learning algorithm that incurs $\widetilde{O}(\sqrt{T})$ regret in expectation over a horizon of length $T$; this guarantee is optimal, up to polylogarithmic factors. We then provide a devaluation-free modification of this algorithm, and show that this results in only a constant factor loss in regret.