Search arXivSearch

arXiv · 2603.25138

Reinforcement learning for quantum processes with memory

Abstract

In reinforcement learning, an agent interacts sequentially with an environment to maximize a reward, receiving only partial, probabilistic feedback. This creates a fundamental exploration-exploitation trade-off: the agent must explore to learn the hidden dynamics while exploiting this knowledge to maximize its target objective. While extensively studied classically, applying this framework to quantum systems requires dealing with hidden quantum states that evolve via unknown dynamics. We formalize this problem via a framework where the environment maintains a hidden quantum memory evolving via unknown quantum channels, and the agent intervenes sequentially using quantum instruments. For this setting, we adapt an optimistic maximum-likelihood estimation algorithm. We extend the analysis to continuous action spaces, allowing us to model general positive operator-valued measures (POVMs). By controlling the propagation of estimation errors through quantum channels and instruments, we prove that the cumulative regret of our strategy scales as $\widetilde{\mathcal{O}}(\sqrt{K})$ over $K$ episodes. Furthermore, via a reduction to the multi-armed quantum bandit problem, we establish information-theoretic lower bounds demonstrating that this sublinear scaling is strictly optimal up to polylogarithmic factors. As a physical application, we consider state-agnostic work extraction. When extracting free energy from a sequence of non-i.i.d. quantum states correlated by a hidden memory, any lack of knowledge about the source leads to thermodynamic dissipation. In our setting, the mathematical regret exactly quantifies this cumulative dissipation. Using our adaptive algorithm, the agent uses past energy outcomes to improve its extraction protocol on the fly, achieving sublinear cumulative dissipation, and, consequently, an asymptotically zero dissipation rate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Josep Lumbreras, Ruo Cheng Huang, Yanglin Hu, Marco Fanizza, Mile Gu. 2026-03-26. Reinforcement learning for quantum processes with memory. https://arxiv.org/abs/2603.25138

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantum Authenticated Key Expansion with Key Recycling

Data privacy and authentication are two main security requirements for remote access and cloud services. While QKD has been explored to address data privacy concerns, oftentimes its use is separate from the client authentication protocol despite implicitly providing authentication. Here, we present a quantum authentication key expansion (QAKE) protocol that (1) integrates both authentication and key expansion within a single protocol, and (2) provides key recycling property - allowing all authentication keys to be reused. We analyse the security of the protocol in a QAKE framework adapted from a classical authentication key exchange (AKE) framework, providing separate security conditions for authentication and data privacy. We experimentally implemented the protocol with appropriate post-selection. Additional results on the security of pseudorandom basis generation in QAKE and decoy state BB84 are provided.

quant-ph

Entanglement as Difference: Reduction-induced Minimal Partial Entropy Difference

Bipartite mixed-state quantum entanglement (QE) and its measures play a crucial role in both theoretical research and practical quantum applications. Its internal structure is far more complex and less well understood compared with bipartite pure-state QE. Some existing measures involve inherently intractable global optimizations, while others are only applicable to highly limited-dimensional quantum systems. Here based on the inherent feature that bipartite QE systems nonseparable necessarily implies that local reduced density matrix differs from its \textquotedblleft native\textquotedblright density matrix, we propose a more physical and intuitive measure termed Reduction-induced Minimal Partial Entropy Difference to quantify arbitrary bipartite mixed-state QE. Partial Von Neumann Entropy is only a pure-state special case of this method. This measure offers intrinsic structural %perspective insights into bipartite QE characterization, thereby establishing itself as a valuable complementary measure. Its intuitive and clear physical picture, combined with relatively low computational complexity and wide applicability, facilitates exploring its potential quantum information applications, hence its conceptual framework and line of thought deserve to be further developed to describe and quantify multipartite QE in the future.

quant-ph

Non-local mass superpositions and optical clock interferometry in atomic ensemble quantum networks

Quantum networks are emerging as powerful platforms for sensing, communication, and fundamental tests of physics. We propose a programmable quantum sensing network based on entangled atomic ensembles, where optical clock qubits realize mass superpositions arising via mass-energy equivalence, as in atom and atom-clock interferometry. Our approach uniquely combines scalability to large atom numbers with minimal control requirements, relying only on collective addressing of internal atomic states. This enables the creation of both non-local and local superpositions with spatial separations beyond those achievable in conventional matter-wave interferometry with single atoms. Starting from Bell-type seed states distributed via photonic channels, collective operations within atomic ensembles coherently build many-body mass superpositions sensitive to gravitational redshift. The resulting architecture implements a non-local Ramsey interferometer, where gravitationally induced phase shifts are imprinted on non-local entangled states and are read out through local measurements at the network nodes. Beyond extending the spatial reach of mass superpositions, our scheme establishes a scalable, programmable platform to probe the interface of quantum mechanics and gravity, and offers a new experimental pathway to test atom and atom-clock interferometer proposals, e.g. for probing gravitational dephasing, in a network-based quantum laboratory.

quant-ph