Search arXiv⌕ Search

arXiv · 2610.00945

TATVA: A Reinforcement Learning Framework for Quantum Circuit Synthesis

Abstract

One of the most important initial steps in quantum computing is high-fidelity quantum state preparation, because errors in it can affect the final computational result. Therefore, one of the major challenges is to design a quantum circuit that can achieve the desired target state accurately with a smaller number of gates. Traditionally, circuits are designed using predefined rules and mathematical decompositions, but this makes it difficult to design circuits for different and complex target states. This paper presents TATVA, a reinforcement learning (RL) system for synthesizing quantum circuits. TATVA views circuit synthesis as a sequential decision problem in which the agents select one gate at a time and calculate the fidelity after each applied gate, with the gate producing the highest fidelity selected for the circuit. The system introduces a parallel architecture that uses two RL agents, Deep Q-Network (DQN) and Proximal Policy Optimization (PPO), along with Qiskit's statevector simulator. The system can synthesize up to 5 qubits while achieving a fidelity of 0.999999 and producing compact circuits. After achieving the desired fidelity, it performs a post-hoc optimization step that reduces the circuit depth while preserving the achieved fidelity. The circuit is finalized using fidelity, success rate, gate count, and circuit depth. Target states that are not used during training are used to test whether the system has learned effectively and can generalize beyond the training states. This approach aims to enable automatic quantum circuit synthesis with high fidelity using reinforcement learning.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Om Bhamare, Aryan Jadhav, Manas Shinde, Prithvi Shinde. 2026-10-01. TATVA: A Reinforcement Learning Framework for Quantum Circuit Synthesis. https://arxiv.org/abs/2610.00945

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Quantum squeezing cannot beat the standard quantum limit

Quantum entanglement between particles is expected to allow one to perform tasks that would otherwise be impossible. In quantum sensing and metrology, entanglement is often claimed to enable a measurement precision that cannot be attained with the same number of particles and time, forgoing entanglement. Two distinct approaches exist: creation of entangled states that either i) respond quicker to the signal, or ii) are associated with lower noise and uncertainty. The second class of states are generally called squeezed states. Here we show that if our definition of success is a precision that is impossible to achieve using the same resources but without entanglement then squeezed states cannot succeed. In doing so we show that a single non-separable squeezed state provides fundamentally no better precision, per unit time, than a single particle.

quant-ph↗

Quantum algorithms for general nonlinear dynamics based on the Carleman embedding

Important nonlinear dynamics, such as those found in plasma and fluid systems, are typically hard to simulate on classical computers. Thus, if fault-tolerant quantum computers could efficiently solve such nonlinear problems, it would be a transformative change for many industries. In a recent breakthrough [Liu et al., PNAS 2021], the first efficient quantum algorithm for solving nonlinear differential equations was constructed, based on a single condition $R<1$, where $R$ characterizes the ratio of nonlinearity to dissipation. This result, however, is limited to the class of purely dissipative systems with negative log-norm, which excludes application to many important problems. In this work, we correct technical issues with this and other prior analysis, and substantially extend the scope of nonlinear dynamical systems that can be efficiently simulated on a quantum computer in a number of ways. Firstly, we extend the existing results from purely dissipative systems to a much broader class of stable systems, and show that every quadratic Lyapunov function for the linearized system corresponds to an independent $R$-number criterion for the convergence of the Carlemen scheme. Secondly, we extend our stable system results to physically relevant settings where conserved polynomial quantities exist. Finally, we provide extensive results for the class of non-resonant systems. With this, we are able to show that efficient quantum algorithms exist for a much wider class of nonlinear systems than previously known, and prove the BQP-completeness of nonlinear oscillator problems of exponential size. In our analysis, we also obtain several results related to the Poincaré-Dulac theorem and diagonalization of the Carleman matrix, which could be of independent interest.

quant-ph↗

Sampled-Based Guided Quantum Walk: Non-variational quantum algorithm for combinatorial optimization

We introduce SamBa-GQW, a novel quantum algorithm for solving binary combinatorial optimization problems of arbitrary degree with no use of any classical optimizer. The algorithm is based on a continuous-time quantum walk on the solution space represented as a graph. The walker explores the solution space to find its way to vertices that minimize the cost function of the optimization problem. The key novelty of our algorithm is an offline classical sampling protocol that gives information about the spectrum of the problem Hamiltonian. Then, the extracted information is used to guide the walker to high quality solutions via a quantum walk with a time-dependent hopping rate. We investigate the performance of SamBa-GQW on several quadratic problems, namely MaxCut, maximum independent set, portfolio optimization, and higher-order polynomial problems such as LABS, MAX-$k$-SAT and a quartic reformulation of the travelling salesperson problem. We empirically demonstrate that SamBa-GQW finds high quality approximate solutions on problems up to a size of $n=30$ qubits by only sampling poly($n$) states among $2^n$ possible decisions. Furthermore, SamBa-GQW compares in par with classically optimized variational approaches, such as the variational guided quantum walk and QAOA (the latter when run on deep circuits). This places Samba-GQW as a promising heuristic to tackle combinatorial problems beyond the NISQ regime.

quant-ph↗