Search arXivSearch

arXiv subjects

Zhan Yu

Publications and source records attributed to Zhan Yu.

At least 19 recordsLinked to original sources

Quantum Quasi-Monte Carlo: a window for pre-asymptotic quantum advantage

Numerical integration with Monte Carlo methods is a central computational task in many scientific and industrial applications, including financial derivative pricing and risk management. Classical Monte Carlo algorithms are computationally demanding: achieving an accuracy $\epsilon$ typically requires a number of function evaluations scaling as $O(1/\epsilon^2)$. Quantum-accelerated Monte Carlo methods based on quantum amplitude estimation can in principle quadratically improve this dependence. However, \textit{quasi}-Monte Carlo methods have not been explored in the quantum context. In this work, we introduce a quantum quasi-Monte Carlo algorithm that combines low-discrepancy nets with quantum amplitude estimation. The proposed method prepares the quasi-random point set coherently in superposition. The method does not yield an asymptotic improvement over classical quasi-Monte Carlo, since the total error separates into a discretization error, determined by the finite net, and a quantum estimation error. Instead, we explore a pre-asymptotic advantage window: for a target accuracy that would classically require $2^q$ low discrepancy points, one can prepare a higher-resolution net of size $2^Q$, with $Q>q$, in superposition and reach the same accuracy using significantly fewer function queries. This window can be controlled by tuning the circuit resolution and amplitude-estimation parameters, making the approach relevant for practical regimes where the number of queries is finite rather than asymptotically large.

quant-ph

Learning Arbitrary Lindbladians from Time Evolution

We study the problem of learning an unknown Markovian open-system generator from access to its physical time evolution. This generator, called a Lindbladian, contains Hamiltonian and dissipative coefficients indexed by an exponentially large family of possible Pauli terms. We propose an efficient algorithm that learns arbitrary Lindbladians from time evolution under minimal assumptions. For a Lindbladian of dynamical strength at most $\Lambda$, the algorithm estimates every coefficient to error $\epsilon$ using $\widetilde O(\Lambda^2/\epsilon^2)$ experiments and $\widetilde O(\Lambda/\epsilon^2)$ total evolution time, together with polynomial classical running time. The algorithm consists of two nonadaptive, ancilla-free, and control-free stages: 1. The support-learning stage outputs a candidate support of size $\mathrm{poly}(\Lambda/\eta)$ that contains every Hamiltonian and dissipative coordinate of magnitude at least $\eta$, using $\widetilde O(\Lambda^2/\eta^2)$ experiments with preparations of product Pauli eigenstates and single-qubit Pauli measurements. 2.The coefficient-learning stage estimates all coefficients in any candidate support of size $M$ to error $\epsilon$, using $\widetilde O(\Lambda^2\log M/\epsilon^{2})$ experiments with preparations of random stabilizer states and measurements in random Clifford bases. Composing the two stages identifies and estimates every coefficient of an arbitrary Lindbladian in polynomial time. The experiment-count and total-evolution-time scalings match the lower bounds up to logarithmic factors, so the algorithm is nearly optimal for learning arbitrary Lindbladians.

quant-ph

On estimating operator norm distance, with optimal trace distance estimation when one state is pure

We investigate the computational complexity of estimating the operator norm distance ${\rm T}_{\infty}(\rho_0,\rho_1)$, defined via the operator norm $\|A\|_{\infty} = \sigma_{\max}(A)$, given ${\rm poly}(n)$-size state-preparation circuits of $n$-qubit quantum states $\rho_0$ and $\rho_1$. We provide efficient quantum estimators for the operator norm distance whose complexity is independent of the rank (and thus the dimension) of the states: 1. When one state is pure, we establish an optimal quantum estimator using $\Theta(1/\epsilon)$ queries to the state-preparation circuits. Consequently, for constant additive error, say $\epsilon=1/5$, our estimator runs in ${\rm poly}(n)$ time. Since the operator norm distance ${\rm T}_{\infty}(|\psi\rangle\!\langle\psi|,\rho)$ is exactly half of the trace distance ${\rm T}(|\psi\rangle\!\langle\psi|,\rho)$, our result also gives rank-independent query complexity for estimating both quantities, whereas the approaches due to van Apeldoorn, Cornelissen, Gily{\'{e}}n, and Nannicini (SODA 2023) and Wang and Zhang (TIT 2024) have query complexity scaling at least linearly with ${\rm rank}(\rho)$, which can be $\exp(n)$ in general. 2. For general quantum states, we also provide a quantum estimator using $\widetilde{O}(1/\epsilon^{3/2})$ queries to the state-preparation circuits, which shows that the corresponding promise problem is ${\sf BQP}$-complete and improves the ${\sf QMA}$ upper bound sketched by Liu and Wang (ESA 2025). Together with an $\Omega(1/\epsilon)$ quantum query complexity lower bound, this leaves only square-root room for improvement. The key intuition behind our estimators is that, when one state is pure, the pure state $|\psi\rangle$ has overlap at least $1/2$ with the top unit eigenvector of $|\psi\rangle\!\langle\psi|-\rho$, reflecting a structural feature specific to the operator norm distance.

quant-ph

Near-Optimal Learning of Local Lindbladians

We study the problem of learning local Lindbladians from black-box access to the physical evolution, where the goal is to estimate all Hamiltonian and dissipative coefficients. For local Lindbladians with low-intersection dissipators and local dynamical strength at most $\Lambda$, we design an algorithm that learns the coefficients to target accuracy $\varepsilon$, whose number of channel uses and total evolution time scale as $\widetilde{O}(\Lambda^2/\varepsilon^2)$ and $\widetilde{O}(\Lambda/\varepsilon^2)$, respectively, with only logarithmic dependence on the number of qubits. The algorithm is built directly from finite-time channel probes. It runs the unknown evolution for short times, estimates the corresponding Pauli transfer matrices from classical shadows, and converts these estimates into Lindbladian coefficients by stable local Fourier inversions. The algorithm is non-adaptive, uses no ancillas, and requires only random product states as inputs followed by random Pauli measurements. The method does not require knowing the structure of the Lindbladian in advance. We prove matching lower bounds that establish the near-optimality of the algorithm in both resources. By constructing a family of single-qubit dephasing Lindbladians, we show that any algorithm, even an adaptive one with arbitrary ancillas and measurements, requires $\Omega(\Lambda^2/\varepsilon^2)$ channel uses and $\Omega(\Lambda/\varepsilon^2)$ total evolution time. In particular, the lower bounds imply that the Heisenberg-limited scaling achievable for Hamiltonian learning is information-theoretically impossible once dissipative coefficients must be estimated.

quant-ph

Generic Frameworks for Distributed Functional Optimization and Learning over Time-Varying Networks

In this paper, we establish a distributed functional optimization (DFO) theory over time-varying networks. The vast majority of existing distributed optimization theories are developed based on Euclidean decision variables. However, for many scenarios in machine learning and statistical learning, such as reproducing kernel spaces or probability measure spaces that use functions or probability measures as fundamental variables, the development of existing distributed optimization theories exhibit obvious theoretical and technical deficiencies. This paper addresses these issues by developing a novel general DFO theory on Banach spaces, allowing functional learning problems in the aforementioned scenarios to be incorporated into our framework for resolution. We study both convex and nonconvex DFO problems and rigorously establish a comprehensive convergence theory of distributed functional mirror descent and distributed functional gradient descent algorithm to solve them. Satisfactory convergence rates are fully derived. The work has provided generic analyzing frameworks for DFO. The established theory is shown to have crucial application value in the kernel-based distributed learning theory over networks.

math.OC

Distributed Stochastic Optimization under Heavy-Tailed Noise: A Federated Mirror Descent Approach with High Probability Convergence

We study the distributed stochastic optimization (DSO) problem under a heavy-tailed noise condition by utilizing a multi-agent system. Despite the extensive research on DSO algorithms used to solve DSO problems under light-tailed noise conditions (such as Gaussian noise), there is a significant lack of study of DSO algorithms in the context of heavy-tailed random noise. Classical DSO approaches in a heavy-tailed setting may present poor convergence behaviors. Therefore, developing DSO methods in the context of heavy-tailed noises is of importance. This work follows this path and we consider the setting that the gradient noises associated with each agent can be heavy-tailed, potentially having unbounded variance. We propose a clipped federated stochastic mirror descent algorithm to solve the DSO problem. We rigorously present a convergence theory and show that, under appropriate rules on the stepsize and the clipping parameter associated with the local noisy gradient influenced by the heavy-tailed noise, the algorithm is able to achieve satisfactory high probability convergence.

math.OC

High-Probability Convergence Theory for Distributed Composite Optimization with Sub-Weibull Noises

With the rapid development of distributed optimization (DO) theory, the distributed stochastic gradient methods (DSGMs) occupy an important position. Although the theory of different DSGMs has been widely established, the main-stream results of existing work are still derived under the condition of light-tailed stochastic gradient noises. Increasing examples from various fields, indicate that, the light-tailed noise model is overly idealized in many practical instances, failing to capture the complexity and variability of noises in real-world scenarios, such as the presence of outliers or extreme values from data science and statistical learning. To address this issue, we propose a new DO framework that incorporates stochastic gradients under sub-Weibull randomness. We study a distributed composite stochastic mirror descent scheme with sub-Weibull gradient noise (DCSMD-SW) for solving a convex distributed composite optimization (DCO) problem over the time-varying multi-agent network. By investigating sub-Weibull randomness in DCSMD for the first time, we show that the algorithm is applicable in some common heavier-tailed noise environments while also guaranteeing good convergence properties. We comprehensively study the convergence performance of DCSMD-SW. Satisfactory high-probability convergence rates are derived for DCSMD-SW without any smoothness requirement. The work also offers a unified analytical framework for several critical cases of both algorithms and noise environments.

math.OC

Theory of Decentralized Robust Kernel-Based Learning

We propose a new decentralized robust kernel-based learning algorithm within the framework of reproducing kernel Hilbert spaces (RKHSs) by utilizing a networked system that can be represented as a connected graph. The robust loss function $\huaL_\sigma$ induced by a windowing function $W$ and a robustness scaling parameter $\sigma>0$ can encompass a broad spectrum of robust losses. Consequently, the proposed algorithm effectively provides a unified decentralized learning framework for robust regression, which fundamentally differs from the existing distributed robust kernel-based learning schemes, all of which are divide-and-conquer based. We rigorously establish a learning theory and offer comprehensive convergence analysis for the algorithm. We show each local robust estimator generated from the decentralized algorithm can be utilized to approximate the regression function. Based on kernel-based integral operator techniques, we derive general high confidence convergence bounds for the local approximating sequence in terms of the mean square distance, RKHS norm, and generalization error, respectively. Moreover, we provide rigorous selection rules for local sample size and show that, under properly selected step size and scaling parameter $\sigma$, the decentralized robust algorithm can achieve optimal learning rates (up to logarithmic factors) in both norms. The parameter $\sigma$ is shown to be essential for enhancing robustness and ensuring favorable convergence behavior. The intrinsic connection among decentralization, sample selection, robustness of the algorithm, and its convergence is clearly reflected.

cs.LG

Simultaneous Estimation of Nonlinear Functionals of a Quantum State

We consider a fundamental task in quantum information theory, estimating the values of $\operatorname{tr}(O\rho)$, $\operatorname{tr}(O\rho^2)$, ..., $\operatorname{tr}(O\rho^k)$ for an observable $O$ and a quantum state $\rho$. We show that $\widetilde\Theta(k)$ samples of $\rho$ are sufficient and necessary to simultaneously estimate all the $k$ values. This means that estimating all the $k$ values is almost as easy as estimating only one of them, $\operatorname{tr}(O\rho^k)$. As an application, our approach advances the sample complexity of entanglement spectroscopy and the virtual cooling for quantum many-body systems. Moreover, we extend our approach to estimating general functionals by polynomial approximation.

quant-ph

Exploring experimental limit of deep quantum signal processing using a trapped-ion simulator

Quantum signal processing (QSP), which enables systematic polynomial transformations on quantum data through sequences of qubit rotations, has emerged as a fundamental building block for quantum algorithms and data re-uploading quantum neural networks. While recent experiments have demonstrated the feasibility of shallow QSP circuits, the inherent limitations in scaling QSP to achieve complex transformations on quantum hardware remain an open and critical question. Here we report the first experimental realization of deep QSP circuits in a trapped-ion quantum simulator. By manipulating the qubit encoded in a trapped $^{43}\textrm{Ca}^{+}$ ion, we demonstrate high-precision simulation of some prominent functions used in quantum algorithms and machine learning, with circuit depths ranging from 15 to 360 layers and implementation time significantly longer than coherence time of the qubit. Our results reveal a crucial trade-off between the precision of function simulation and the concomitant accumulation of hardware noise, highlighting the importance of striking a balance between circuit depth and accuracy in practical QSP implementation. This work addresses a key gap in understanding the scalability and limitations of QSP-based algorithms on quantum hardware, providing valuable insights for developing quantum algorithms as well as practically realizing quantum singular value transformation and data re-uploading quantum machine learning models.

quant-ph

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers

This tutorial intends to introduce readers with a background in AI to quantum machine learning (QML) -- a rapidly evolving field that seeks to leverage the power of quantum computers to reshape the landscape of machine learning. For self-consistency, this tutorial covers foundational principles, representative QML algorithms, their potential applications, and critical aspects such as trainability, generalization, and computational complexity. In addition, practical code demonstrations are provided in https://qml-tutorial.github.io/ to illustrate real-world implementations and facilitate hands-on learning. Together, these elements offer readers a comprehensive overview of the latest advancements in QML. By bridging the gap between classical machine learning and quantum computing, this tutorial serves as a valuable resource for those looking to engage with QML and explore the forefront of AI in the quantum era.

quant-ph

Amortized Stabilizer R\'enyi Entropy of Quantum Dynamics

Unraveling the secrets of how much nonstabilizerness a quantum dynamic can generate is crucial for harnessing the power of magic states, the essential resources for achieving quantum advantage and realizing fault-tolerant quantum computation. In this work, we introduce the amortized $\alpha$-stabilizer R\'enyi entropy, a magic monotone for unitary operations that quantifies the nonstabilizerness generation capability of quantum dynamics. Amortization is key in quantifying the magic of quantum dynamics, as we reveal that nonstabilizerness generation can be enhanced by prior nonstabilizerness in input states when considering the $\alpha$-stabilizer R\'enyi entropy, while this is not the case for robustness of magic or stabilizer extent. We demonstrate the versatility of the amortized $\alpha$-stabilizer R\'enyi entropy in investigating the nonstabilizerness resources of quantum dynamics of computational and fundamental interest. In particular, we establish improved lower bounds on the $T$-count of quantum Fourier transforms and the quantum evolutions of one-dimensional Heisenberg Hamiltonians, showcasing the power of this tool in studying quantum advantages and the corresponding cost in fault-tolerant quantum computation.

quant-ph

Quantum Transformer: Accelerating model inference via quantum linear algebra

Powerful generative artificial intelligence from large language models (LLMs) harnesses extensive computational resources for inference. In this work, we investigate the transformer architecture, a key component of these models, under the lens of fault-tolerant quantum computing. We develop quantum subroutines to construct the building blocks in the transformer, including the self-attention, residual connection with layer normalization, and feed-forward network. As an important subroutine, we show how to efficiently implement the Hadamard product and element-wise functions of matrices on quantum computers. Our algorithm prepares an amplitude encoding of the transformer output, which can be measured for prediction or use in the next layer. We find that the matrix norm of the input sequence plays a dominant role in the quantum complexity. With numerical experiments on open-source LLMs, including for bio-informatics applications, we demonstrate the potential of a quantum speedup for transformer inference in practical regimes.

quant-ph

Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits

Parameterized quantum circuits (PQCs) have emerged as a promising approach for quantum neural networks. However, understanding their expressive power in accomplishing machine learning tasks remains a crucial question. This paper investigates the expressivity of PQCs for approximating general multivariate function classes. Unlike previous Universal Approximation Theorems for PQCs, which are either nonconstructive or rely on parameterized classical data processing, we explicitly construct data re-uploading PQCs for approximating multivariate polynomials and smooth functions. We establish the first non-asymptotic approximation error bounds for these functions in terms of the number of qubits, quantum circuit depth, and number of trainable parameters. Notably, we demonstrate that for approximating functions that satisfy specific smoothness criteria, the quantum circuit size and number of trainable parameters of our proposed PQCs can be smaller than those of deep ReLU neural networks. We further validate the approximation capability of PQCs through numerical experiments. Our results provide a theoretical foundation for designing practical PQCs and quantum neural networks for machine learning tasks that can be implemented on near-term quantum devices, paving the way for the advancement of quantum machine learning.

quant-ph

Learning Theory of Distribution Regression with Neural Networks

In this paper, we aim at establishing an approximation theory and a learning theory of distribution regression via a fully connected neural network (FNN). In contrast to the classical regression methods, the input variables of distribution regression are probability measures. Then we often need to perform a second-stage sampling process to approximate the actual information of the distribution. On the other hand, the classical neural network structure requires the input variable to be a vector. When the input samples are probability distributions, the traditional deep neural network method cannot be directly used and the difficulty arises for distribution regression. A well-defined neural network structure for distribution inputs is intensively desirable. There is no mathematical model and theoretical analysis on neural network realization of distribution regression. To overcome technical difficulties and address this issue, we establish a novel fully connected neural network framework to realize an approximation theory of functionals defined on the space of Borel probability measures. Furthermore, based on the established functional approximation results, in the hypothesis space induced by the novel FNN structure with distribution inputs, almost optimal learning rates for the proposed distribution regression model up to logarithmic terms are derived via a novel two-stage error decomposition technique.

stat.ML

Distributed Gradient Descent for Functional Learning

In recent years, different types of distributed and parallel learning schemes have received increasing attention for their strong advantages in handling large-scale data information. In the information era, to face the big data challenges {that} stem from functional data analysis very recently, we propose a novel distributed gradient descent functional learning (DGDFL) algorithm to tackle functional data across numerous local machines (processors) in the framework of reproducing kernel Hilbert space. Based on integral operator approaches, we provide the first theoretical understanding of the DGDFL algorithm in many different aspects of the literature. On the way of understanding DGDFL, firstly, a data-based gradient descent functional learning (GDFL) algorithm associated with a single-machine model is proposed and comprehensively studied. Under mild conditions, confidence-based optimal learning rates of DGDFL are obtained without the saturation boundary on the regularity index suffered in previous works in functional regression. We further provide a semi-supervised DGDFL approach to weaken the restriction on the maximal number of local machines to ensure optimal rates. To our best knowledge, the DGDFL provides the first divide-and-conquer iterative training approach to functional learning based on data samples of intrinsically infinite-dimensional random functions (functional covariates) and enriches the methodologies for functional data analysis.

stat.ML

Efficient information recovery from Pauli noise via classical shadow

The rapid advancement of quantum computing has led to an extensive demand for effective techniques to extract classical information from quantum systems, particularly in fields like quantum machine learning and quantum chemistry. However, quantum systems are inherently susceptible to noises, which adversely corrupt the information encoded in quantum systems. In this work, we introduce an efficient algorithm that can recover information from quantum states under Pauli noise. The core idea is to learn the necessary information of the unknown Pauli channel by post-processing the classical shadows of the channel. For a local and bounded-degree observable, only partial knowledge of the channel is required rather than its complete classical description to recover the ideal information, resulting in a polynomial-time algorithm. This contrasts with conventional methods such as probabilistic error cancellation, which requires the full information of the channel and exhibits exponential scaling with the number of qubits. We also prove that this scalable method is optimal on the sample complexity and generalise the algorithm to the weight contracting channel. Furthermore, we demonstrate the validity of the algorithm on the 1D anisotropic Heisenberg-type model via numerical simulations. As a notable application, our method can be severed as a sample-efficient error mitigation scheme for Clifford circuits.

quant-ph

Quantum Phase Processing and its Applications in Estimating Phase and Entropies

Quantum computing can provide speedups in solving many problems as the evolution of a quantum system is described by a unitary operator in an exponentially large Hilbert space. Such unitary operators change the phase of their eigenstates and make quantum algorithms fundamentally different from their classical counterparts. Based on this unique principle of quantum computing, we develop a new algorithmic toolbox "quantum phase processing" that can directly apply arbitrary trigonometric transformations to eigenphases of a unitary operator. The quantum phase processing circuit is constructed simply, consisting of single-qubit rotations and controlled-unitaries, typically using only one ancilla qubit. Besides the capability of phase transformation, quantum phase processing in particular can extract the eigen-information of quantum systems by simply measuring the ancilla qubit, making it naturally compatible with indirect measurement. Quantum phase processing complements another powerful framework known as quantum singular value transformation and leads to more intuitive and efficient quantum algorithms for solving problems that are particularly phase-related. As a notable application, we propose a new quantum phase estimation algorithm without quantum Fourier transform, which requires the fewest ancilla qubits and matches the best performance so far. We further exploit the power of our method by investigating a plethora of applications in Hamiltonian simulation, entanglement spectroscopy and quantum entropies estimation, demonstrating improvements or optimality for almost all cases.

quant-ph