Search arXiv⌕ Search

arXiv · 2610.02745

AI-Assisted GPU optimization of the Stochastic Variational Method for Few-Body Boson System

Abstract

The stochastic variational method (SVM) is one of the most powerful methods to solve quantum few-body systems precisely in various fields, such as nuclear and atomic physics. To the best of our knowledge, no SVM code optimized for GPUs has been reported. We developed a parallel SVM code for few-body clusters of $^4$He atoms and optimized it for two GPU architectures, NVIDIA GH200 and AMD Instinct MI300A, with the entire code written by an AI coding agent. On the algorithmic side, we introduce the secular-equation method with the Gu--Eisenstat prescription for solving the generalized eigenvalue problem within the SVM framework. The main part of the code tuning is to batch the many small matrices so that the GPUs are used efficiently, together with an array layout in memory optimized for coalesced addressing. For the five-body system with 4800 SVM basis states, the tuned SVM code runs 16.2 (14.7) times faster on the NVIDIA GH200 (AMD MI300A) than the same tuned code on the Intel Xeon Max CPU system, and 168 (153) times faster than the first working CPU implementation. The achieved performance brings 6-body and larger cluster systems within reach.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Noritaka Shimizu, Shigeyoshi Aoyama, Yusuke Tsunoda, Akira Nukada. 2026-10-02. AI-Assisted GPU optimization of the Stochastic Variational Method for Few-Body Boson System. https://arxiv.org/abs/2610.02745

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Learning Lattice Parameters from Powder X-Ray Diffraction Data Using Invariants

We present a machine learning (ML) method to determine unit cell parameters from powder X-Ray diffraction (XRD) data using a novel invariant lattice representation. In ML, the data representation used can have a substantial impact on the prediction quality. Previous approaches have directly predicted lattice parameters ($a,b,c,α,β,γ$) from XRD inputs. However, these parameters depend strongly on the unit cell reduction or convention used. In this work, we construct an invariant representation of the reciprocal lattice that is independent of primitive cell convention, based on the bispectrum--a descriptor built from spherical harmonic projections of lattice points. The calculation of the lattice bispectrum is differentiable, and we demonstrate how to invert it using gradient-based optimization initialized from a nearest-neighbor lookup. We show that with a fixed ML model architecture, using the lattice bispectrum as the ML target rather than the unit cell parameters leads to more accurate lattice parameter predictions. For example, using the MP-20 dataset, the bispectrum reduces length mean absolute percentage error (MAPE) from 11.08% to 2.43% and angle MAPE from 12.15% to 2.45% compared to direct prediction with the same model architecture. We additionally benchmark our approach against pre-existing XRD to crystal structure models such as Crystalyze and assess its performance on experimental data. Beyond unit cell representation, we anticipate this invariant lattice representation could serve more broadly as a geometry-aware target for other crystallographic machine learning tasks such as structure generation.

physics.comp-ph↗

Mechanism-resolved phase-field fracture of composite shells with a certified admissible constitutive operator

Phase-field models of fracture in fibre-reinforced shells are usually formulated against a single closed-form stored energy, so that the stress, consistent tangent and plane-stress condensation are derived and verified for that expression alone. Here the constitutive law enters instead as a replaceable operator. The Green-Lagrange strain of a geometrically exact Reissner-Mindlin shell is pulled back through the Cholesky factor of the reference metric and resolved on four mutually orthogonal invariants with an exact closure identity, making the separation into fibre, inter-fibre and interaction channels a change of basis rather than a modelling assumption. On a curved midsurface, the metric, Cholesky factor and fibre direction vary through the thickness through the shifter I - zeta b, and the invariants inherit that dependence. The tension-compression split, channel degradation, plane-stress condensation, consistent tangent and finite-element assembly are all expressed in the channel potentials and their derivatives, allowing either closed-form or learned operators. The fibre-transverse interaction energy is stored once and degraded by the product of the fields whose mechanisms it couples. Six conditions define admissible operators; the two linear invariants are proved convex in the deformation gradient, the geometric tangent is shown independent of the material tangent, and the plane-stress condensation is regular wherever the residual stiffness is positive. A three-tier protocol certifies the algebra, solver state and admission of a state as training data. Numerical studies on notched IM7/8552 shells verify the discretisation in thin and curved regimes and show that fibre orientation governs the damage envelope and load capacity, while curvature drives a through-thickness fracture asymmetry that a single midsurface field cannot represent.

physics.comp-ph↗

HFS-TransNet: A Hybrid Fixed-Stress Transferable Neural Network for Quasi-Static Biot Poroelasticity

The quasi-static Biot system, which governs coupled fluid-solid interactions in poroelastic media, poses significant computational challenges due to its strong coupling and multiscale nature. To address these, we propose a hybrid fixed-stress transferable neural network (HFS-TransNet) method. In HFS-TransNet, a data-driven TransNet module first approximates observed data on an initialization domain to supply a high-quality starting point for the subsequent physics-informed stage, which then drives fixed-stress splitting iterations via TransNet and admits an error bound. After each iteration, a greedy update strategy on a validation domain prevents model degradation and a relative tolerance criterion governs convergence termination. HFS-TransNet thereby inherits the robustness of the fixed-stress splitting scheme, the flexibility of data-driven modeling, and the transferability of TransNet, effectively overcoming the strong coupling and locking instability inherent in the Biot system. Ablation studies and comparisons with FS-FEM and FS-PINN confirm HFS-TransNet's superior performance in capturing multiscale coupled physics across varying physical parameters and boundary conditions. By bridging classical decoupled iterative schemes with hybrid scientific machine learning, HFS-TransNet offers a novel perspective for complex poroelastic simulations, as demonstrated by its successful application to brain edema simulations.

physics.comp-ph↗