arXiv · 2009.12511
Near-Optimal MNL Bandits Under Risk Criteria
Abstract
We study MNL bandits, which is a variant of the traditional multi-armed bandit problem, under risk criteria. Unlike the ordinary expected revenue, risk criteria are more general goals widely used in industries and bussiness. We design algorithms for a broad class of risk criteria, including but not limited to the well-known conditional value-at-risk, Sharpe ratio and entropy risk, and prove that they suffer a near-optimal regret. As a complement, we also conduct experiments with both synthetic and real data to show the empirical performance of our proposed algorithms.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guangyu Xi, Chao Tao, Yuan Zhou. 2021-03-16. Near-Optimal MNL Bandits Under Risk Criteria. https://arxiv.org/abs/2009.12511
Cite the original work for its findings. Save a collection to share your selection of sources.