Search arXivSearch

arXiv · 2104.14539

Model-Free Quantum Control with Reinforcement Learning

Abstract

Model bias is an inherent limitation of the current dominant approach to optimal quantum control, which relies on a system simulation for optimization of control policies. To overcome this limitation, we propose a circuit-based approach for training a reinforcement learning agent on quantum control tasks in a model-free way. Given a continuously parameterized control circuit, the agent learns its parameters through trial-and-error interaction with the quantum system, using measurement outcomes as the only source of information about the quantum state. Focusing on control of a harmonic oscillator coupled to an ancilla qubit, we show how to reward the learning agent using measurements of experimentally available observables. We train the agent to prepare various non-classical states using both unitary control and control with adaptive measurement-based quantum feedback, and to execute logical gates on encoded qubits. This approach significantly outperforms widely used model-free methods in terms of sample efficiency. Our numerical work is of immediate relevance to superconducting circuits and trapped ions platforms where such training can be implemented in experiment, allowing complete elimination of model bias and the adaptation of quantum control policies to the specific system in which they are deployed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, M. H. Devoret. 2021-12-06. Model-Free Quantum Control with Reinforcement Learning. https://doi.org/10.1103/physrevx.12.011059

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Single-Ensemble Multiparameter Squeezing with Qudits

Conventional spin squeezing enhances a single sensing channel. Here, we show how internal qudit levels enable simultaneous multiparameter squeezing within one ensemble. In two-component magnetometry, a qutrit sensor provides two orthogonal and weakly compatible channels. A collective twisting interaction squeezes both responses while preserving joint attainability of the ultimate sensitivity. The sensing gain is quantified by using a matrix generalization of the Wineland sensitivity that retains both noise correlations and cross-channel response. An interaction-based echo amplifies the signal to overcome noise from a fixed local joint readout, yielding a simulated $13~\mathrm{dB}$ gain over the product-state standard quantum limit for $N=128$ qutrits. More generally, we use the single-site quantum Fisher information matrix to select reference states and channel quadratures for prescribed sensing tasks. The tangent geometry permits at most $d-1$ independent, weakly compatible channels around a common pure reference state for a $d$-level sensor. Our work provides a constructive task-to-protocol map for multiparameter squeezing in a single qudit ensemble.

quant-ph

A Design Space Study of Density Matrix Parameterizations for Diffusion-Based Quantum State Tomography

Diffusion-based quantum state tomography (QST) has shown promising results, but all existing methods implicitly adopt a single parameterization (typically Cholesky) without systematic evaluation. We present the first design space study of density matrix parameterizations for diffusion QST, introducing a geometric framework based on the Jacobian Gram matrix $\mathbf{J}^\top\mathbf{J}$. Our calibration of seven parameterizations at 2- and 3-qubit scales, validated by end-to-end training, reveals that \emph{geometric conditioning alone does not predict end-to-end performance}: at 3-qubit scale, Hermitian direct ($κ= 2.0\times$) performs worse than Cholesky ($κ= 27\times$) at all shot levels---a $13.5\times$ isotropy advantage that translates into a fidelity \emph{disadvantage} of up to $+0.51$. The 2-qubit ranking (Hermitian $>$ Bloch) reverses at 3 qubits (Bloch 0.907 vs.\ Hermitian 0.394). We provide a geometric explanation: unbounded parameterizations suffer projection-induced information loss because the PSD constraint couples diagonal and off-diagonal coordinates in ways the unconstrained model cannot respect, whereas the Bloch representation places the maximally mixed state at the center of the valid region, minimizing projection loss.

quant-ph

Entanglement free Metrology Exploiting Multimode Hong Ou Mandel Sensor Advantage

The Hong-Ou-Mandel (HOM) interference in the multimode frequency domain has been explored for precision metrology, with several experimental demonstrations exploiting its robustness against dispersion and phase noise, as well as its large dynamic range and compatibility with fragile samples. Conventional multimode HOM metrology exploits frequency-entangled states, which naturally satisfy bosonic exchange symmetry under any centered symmetric joint spectral distribution, to provide these advantages. However, these entangled states are typically generated via spontaneous parametric down-conversion (SPDC), requiring strong pump lasers that hinder practical implementation. In this paper, we employ frequency product states, which do not possess entanglement or path-mode exchange symmetry, as the probe state and post-select measurement outcomes exhibiting frequency anti-correlation. Our results demonstrate that these advantages,peak narrowing, dispersion cancellation, phase-noise immunity, a large dynamic range, and compatibility with fragile samples, arise neither from entanglement nor from bosonic exchange symmetry, but rather from spectral anti-correlation. We further show that entanglement is not the source of the measurement precision: the entanglement-free approach attains the same quantum Fisher information as the entangled-state scheme, indicating that the fundamental precision limit does not originate from entanglement.

quant-ph