arXiv · 2203.10214
Thompson Sampling on Asymmetric $\alpha$-Stable Bandits
Abstract
In algorithm optimization in reinforcement learning, how to deal with the exploration-exploitation dilemma is particularly important. Multi-armed bandit problem can optimize the proposed solutions by changing the reward distribution to realize the dynamic balance between exploration and exploitation. Thompson Sampling is a common method for solving multi-armed bandit problem and has been used to explore data that conform to various laws. In this paper, we consider the Thompson Sampling approach for multi-armed bandit problem, in which rewards conform to unknown asymmetric $\alpha$-stable distributions and explore their applications in modelling financial and wireless data.
Explore related subjects
Keep this discovery
Zhendong Shi, Ercan E. Kuruoglu, Xiaoli Wei. 2022-03-19. Thompson Sampling on Asymmetric $\alpha$-Stable Bandits. https://arxiv.org/abs/2203.10214
Cite the original work for its findings. Save a collection to share your selection of sources.