Search arXivSearch

arXiv subjects

Ali Rad

Publications and source records attributed to Ali Rad.

4 recordsLinked to original sources

Rate or Fate? RLV$^\varepsilon$R: Reinforcement Learning with Verifiable Noisy Rewards

Reinforcement learning with verifiable rewards (RLVR) is a simple but powerful paradigm for training LLMs: sample a completion, verify it, and update. In practice, however, the verifier is almost never clean--unit tests probe only limited corner cases; human and synthetic labels are imperfect; and LLM judges (e.g., RLAIF) are noisy and can be exploited--and this problem worsens on harder domains (especially coding) where tests are sparse and increasingly model-generated. We ask a pragmatic question: does the verification noise merely slow down the learning (rate), or can it flip the outcome (fate)? To address this, we develop an analytically tractable multi-armed bandit view of RLVR dynamics, instantiated with GRPO and validated in controlled experiments. Modeling false positives and false negatives and grouping completions into recurring reasoning modes yields a replicator-style (natural-selection) flow on the probability simplex. The dynamics decouples into within-correct-mode competition and a one-dimensional evolution for the mass on incorrect modes, whose drift is determined solely by Youden's index J=TPR-FPR. This yields a sharp phase transition: when J>0, the incorrect mass is driven toward extinction (learning); when J=0, the process is neutral; and when J<0, incorrect modes amplify until they dominate (anti-learning and collapse). In the learning regime J>0, noise primarily rescales convergence time ("rate, not fate"). Experiments on verifiable programming tasks under synthetic noise reproduce the predicted J=0 boundary. Beyond noise, the framework offers a general lens for analyzing RLVR stability, convergence, and algorithmic interventions.

cs.LG

Analog Quantum Simulator of a Quantum Field Theory with Fermion-Spin Systems in Silicon

Simulating fermions coupled to spin degrees of freedom, relevant for a range of quantum field theories, represents a promising application for quantum simulators. Mapping fermions to qubits is challenging in $2+1$ and higher spacetime dimensions, and mapping bosons demands substantial quantum-computational overhead. These features complicate the realization of mixed fermion-boson quantum systems in digital quantum computers. We propose a native fermion-(large-)spin analog quantum simulator by utilizing dopant arrays in silicon. Specifically, we show how to use a dynamical lattice of coupled nuclear spins and conduction-band electrons to realize a quantum field theory: an extended Jackiw-Rebbi model involving coupled fermions and quantum rotors. We demonstrate the feasibility of observing dynamical mass generation and a confinement-deconfinement quantum phase transition in 1+1 dimensions on this platform, even in the presence of strong long-range Coulomb interactions. Furthermore, we employ finite-temperature Hartree-Fock-Bogoliubov simulations to investigate the dynamics of mass generation in two-dimensional square and honeycomb arrays, showing that this phenomenon can be simulated with realistic experimental parameters. Our findings reveal two distinct phases, and demonstrate robustness against the addition of Coulomb interactions. Finally, we discuss experimental signatures of the phases through transport and local charge sensing in dopant arrays. This study lays the foundation for quantum simulations of quantum field theories exhibiting fermions coupled to spin degrees of freedom using donors in silicon.

quant-ph

Deep Quantum Neural Networks are Gaussian Process

The overparameterization of variational quantum circuits, as a model of Quantum Neural Networks (QNN), not only improves their trainability but also serves as a method for evaluating the property of a given ansatz by investigating their kernel behavior in this regime. In this study, we shift our perspective from the traditional viewpoint of training in parameter space into function space by employing the Bayesian inference in the Reproducing Kernel Hilbert Space (RKHS). We observe the influence of initializing parameters using random Haar distribution results in the QNN behaving similarly to a Gaussian Process (QNN-GP) at wide width or, empirically, at a deep depth. This outcome aligns with the behaviors observed in classical neural networks under similar circumstances with Gaussian initialization. Moreover, we present a framework to examine the impact of finite width in the closed-form relationship using a $ 1/d$ expansion, where $d$ represents the dimension of the circuit's Hilbert space. The deviation from Gaussian output can be monitored by introducing new quantum meta-kernels. Furthermore, we elucidate the relationship between GP and its parameter space equivalent, characterized by the Quantum Neural Tangent Kernels (QNTK). This study offers a systematic way to study QNN behavior in over- and under-parameterized scenarios, based on the perturbation method, and addresses the limitations of tracking the gradient descent methods for higher-order corrections like dQNTK and ddQNTK. Additionally, this probabilistic viewpoint lends itself naturally to accommodating noise within our model.

quant-ph

Surviving The Barren Plateau in Variational Quantum Circuits with Bayesian Learning Initialization

Variational quantum-classical hybrid algorithms are seen as a promising strategy for solving practical problems on quantum computers in the near term. While this approach reduces the number of qubits and operations required from the quantum machine, it places a heavy load on a classical optimizer. While often under-appreciated, the latter is a computationally hard task due to the barren plateau phenomenon in parameterized quantum circuits. The absence of guiding features like gradients renders conventional optimization strategies ineffective as the number of qubits increases. Here, we introduce the fast-and-slow algorithm, which uses Bayesian Learning to identify a promising region in parameter space. This is used to initialize a fast local optimizer to find the global optimum point efficiently. We illustrate the effectiveness of this method on the Bars-and-Stripes (BAS) quantum generative model, which has been studied on several quantum hardware platforms. Our results move variational quantum algorithms closer to their envisioned applications in quantum chemistry, combinatorial optimization, and quantum simulation problems.

quant-ph