Search arXivSearch

arXiv · 2604.13096

Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models

Abstract

We propose and analyse a class of analytically solvable models of quantum reinforcement learning (QRL), formulated as finite-horizon Markov decision processes in finite-dimensional Hilbert spaces. The models are built around a `unitary-control-then-measure' protocol, in which a learning agent applies unitary transformations to a quantum state and interleaves each control step with a projective measurement onto a prescribed reference basis. Exact closed-form expressions for trajectory probabilities, rewards, and the expected return are derived for four concrete realisations: a closed-chain and an anti-periodic qubit implementation, a qutrit model with ladder coupling, and a four-level two-qubit system. Two structural features of these QRL protocols are then analysed. First, we identify and quantify the reduction in the computational complexity of the expected return, from the nominally exponential $O(e^N)$ scaling in the trajectory length~$N$ to an explicit power-law $O(N^{\mathcal{I}})$, driven by two rigorously established mechanisms, a trajectory equivalence and a sparsity of the transition graph, besides a third, conjectured one: a spectral concentration of the return, at the optimal policy, onto the polynomially populated trajectory classes. Second, we characterise the degeneracy of optimal policies. The low-dimensional models exhibit unique optima whose asymptotic behaviour with~$N$ is governed by the quantum Zeno effect, while the four-level system displays both plateau-type quasi-degeneracy at large horizons and genuine discrete degeneracy at critical energy parameters -- phenomena with no counterpart in the measurement-free quantum optimal control landscape.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrea Cintio, Alessandro Michelangeli, Dmitrii Tsutskov. 2026-07-02. Complexity scaling and optimal policy degeneracy in quantum reinforcement learning via analytically solvable unitary-control-then-measure models. https://arxiv.org/abs/2604.13096

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Proof of Liu's Conjecture on the Fundamental Triangle Inequality

Let $a,b,c$ be the side lengths of a triangle, and let $R$ and $r$ denote its circumradius and inradius, respectively. Liu proposed the conjecture \[ \sum_{\text{cyc}} \left(\frac{a(b+c-a)}{bc}\right)^k \ge 2+\left(\frac{2r}{R}\right)^k,\qquad k>1, \] with the reverse inequality for $k<1$. We prove this conjecture by reducing it to an algebraic inequality for three positive variables with prescribed sum and product. We also determine the equality cases.

math.GM

A quadratic critical-value conjecture for the fifth Bessel moment

We conjecture an explicit evaluation of the pure fifth Bessel moment $\int_0^\infty K_0(t)^5\,dt$ as a quadratic expression in the critical value $L(f,2)$ of the weight-three, level-60 newform $f$ (LMFDB orbit 60.3.b.a) identified in the twisted fifth-moment modularity theorem of Lim, Tu and Yu, with coefficients in $\mathbb{Q}(\sqrt{5})$ and the square taken before the real and imaginary parts. Directed interval computations, using no stored Bessel or $L$-values, bound the absolute discrepancy by $10^{-358}$. We prove three exact modular identities for $f$: the Petersson-norm formula $\langle f,f\rangle_{60} = \frac{3(5-\sqrt{5})}{2π^4}|L(f,2)|^2$, the coefficient-conjugation relation $L(f^σ,2) = κL(f,2)$ with explicit $κ\in \mathbb{Q}(\sqrt{5},i)$, and the twisted symmetric-square evaluation $L(χ_{-4}\mathrm{Sym}^2 f,2) = \sqrt{15}\,π^2 \langle f,f\rangle_{60}$, together with $L(χ_{-4}\mathrm{Sym}^2 f,3) = π^4\langle f,f\rangle_{60}/8$, in the full Euler-factor normalization of Lim, Tu and Yu. The last identity shows that the companion norm conjecture $D_{5,\mathrm{odd}} = \frac{3\sqrt{15}(5-\sqrt{5})}{2}|L(f,2)|^2$ is equivalent to the symmetric-square conjecture $D_{5,\mathrm{odd}} = π^2 L(χ_{-4}\mathrm{Sym}^2 f,2)$ of Lim, Tu and Yu, while the exact relation $D_{5,\mathrm{even}} = π^2 D_{5,\mathrm{odd}}/(2\sqrt{15})$ follows from Chuang's period formulas. Every Bessel-to-modular equality, including the individual-period formula, remains conjectural. Complete proofs, exact rational certificates and verification programs are included as ancillary files.

math.GM