Search arXivSearch

arXiv · 2507.04438

Quantum Algorithms for Bandits with Knapsacks with Improved Regret and Time Complexities

Abstract

Bandits with knapsacks (BwK) constitute a fundamental model that combines aspects of stochastic integer programming with online learning. Classical algorithms for BwK with a time horizon $T$ achieve a problem-independent regret bound of ${O}(\sqrt{T})$ and a problem-dependent bound of ${O}(\log T)$. In this paper, we initiate the study of the BwK model in the setting of quantum computing, where both reward and resource consumption can be accessed via quantum oracles. We establish both problem-independent and problem-dependent regret bounds for quantum BwK algorithms. For the problem-independent case, we demonstrate that a quantum approach can improve the classical regret bound by a factor of $(1+\sqrt{B/\mathrm{OPT}_\mathrm{LP}})$, where $B$ is budget constraint in BwK and $\mathrm{OPT}_{\mathrm{LP}}$ denotes the optimal value of a linear programming relaxation of the BwK problem. For the problem-dependent setting, we develop a quantum algorithm using an inexact quantum linear programming solver. This algorithm achieves a quadratic improvement in terms of the problem-dependent parameters, as well as a polynomial speedup of time complexity on problem's dimensions compared to classical counterparts. Compared to previous works on quantum algorithms for multi-armed bandits, our study is the first to consider bandit models with resource constraints and hence shed light on operations research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuexin Su, Ziyi Yang, Peiyuan Huang, Tongyang Li, Yinyu Ye. 2025-07-06. Quantum Algorithms for Bandits with Knapsacks with Improved Regret and Time Complexities. https://arxiv.org/abs/2507.04438

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Enhanced measurements on quantum computers via the simultaneous probing of non-commuting Pauli operators

Measuring the state of quantum computers is a highly non-trivial task, with implications for virtually all quantum algorithms. A promising avenue is multi-copy schemes, where identical copies of a quantum state are measured jointly so that all Pauli operators within the considered observable can be simultaneously assessed. Here, we present a first implementation of such a two-copy scheme in a measurement protocol. Based on Bayesian statistics, it accurately estimates not only the average of the desired observable but also the error en route. This enables an adaptive shot-allocation algorithm that preferentially samples the most uncertain Pauli terms. In regimes with many non-commuting Pauli operators, this ``double'' scheme can outperform the state-of-the-art measurement protocol in minimizing total shots for a given precision. We also numerically confirm the finding in previous theoretical works that the two-copy scheme incurs an overhead due to the square-root relationship between the variance of measured quantities and the number of measurement shots.

quant-ph

Thermodynamics of a phaseonium-driven optomechanical Otto engine

We study an optomechanical Otto engine whose working medium is a single-mode cavity driven by beams of coherently prepared three-level phaseonium atoms. The atoms are not thermal reservoirs in the Gibbs sense; rather, their populations and ground-state coherence set the detailed-balance ratio of the cavity collision map, so that the field relaxes to a Gibbs state at an operational apparent temperature. We combine the finite-time collision-model dynamics with radiation-pressure work extraction and compare three reservoir preparations: a thermal reference at the same apparent temperatures, an incoherent atomic beam with the same populations, and the coherent phaseonium beam. We show that the phaseonium isochore charges the cavity passively: the cavity ergotropy and energy-basis coherence remain zero up to numerical precision, while the state converges to the Gibbs fixed point selected by the apparent detailed balance. We further estimate lower bounds on the cost of preparing the atomic populations and coherence, showing that the relevant advantage of phaseonium is a resource-preparation tradeoff rather than a cost-free enhancement over a thermal bath at the same temperature. Finally, we assess the finite-time performance of a two-cavity cascade with additive mechanical work accounting. Over the investigated coherence-phase range, the cascade produces approximately $47\%$--$52\%$ more power than the single-cavity engine while requiring only $65\%$--$68\%$ of the hot and cold phaseonium atoms needed by two independent engines, resulting in a $9\%$--$15\%$ enhancement of power per injected atom over a complete cycle.

quant-ph

Dynamical Correlation of the Post-quench Non-thermal Equilibrium State

After a quantum quench, the integrable system is expected to relax to a non-thermal equilibrium state (NTES) whose local properties are believed to be governed by a generalized Gibbs ensemble (GGE). Combining quench action and the form factor approach, we compute the field-field correlation in the NTES produced by an interaction quench of the Lieb-Liniger model. The spectral distribution is shown to be qualitatively different from that of a thermal equilibrium state (TES): a new dispersion branch appears whose microscopic mechanism can be traced to the algebraic decaying tail for the root density distribution function, and indicates the existence of a broader family of NTES featuring similar spectral property.

quant-ph