Search arXiv⌕ Search

arXiv · 2609.29554

Accelerating Multiple-Precision LU Decomposition with Ozaki Scheme II

Abstract

Solving ill-conditioned linear systems needs LU decomposition in precisions beyond binary64. The standard approaches -- GMP/MPFR, or multi-component arithmetic such as double-double -- leave every scalar multiply-add inside the O(n^3) update in multiple precision, and so cannot exploit the low-precision matrix engines that dominate current hardware. We build a blocked, partially pivoted LU decomposition on Ozaki scheme II, which replaces the multiple-precision GEMM by exact integer modular products followed by an explicit CRT reconstruction, and implement it for multi-component (DD/TD/QD) and arbitrary precision, on CPUs and GPUs. Two contributions make this practical: a direct conversion between the non-overlapping expansion format and the internal fixed-point representation, which removes the MPFR round trip and is worth a factor of 2.4--4.7; and new FP16, FP8 and binary64 GPU back-ends, the FP8 one using a balanced base-17 two-digit encoding that keeps the full 362.8-bit CRT capacity of INT8. On an Arm/GB10 and an x86/H100 the fastest back-end changes with the machine: INT8 wins on GB10, whereas on H100 binary64 is fastest for QD, beating both the native implementation and INT8. The essential point of Ozaki scheme II is thus not to use a low-precision engine but to choose the format maximising bits-per-modulus times engine throughput. For the Lotkin matrix at p ~ 1.2 log2 cond(A) we reach relative errors of 10^-645 at n=2048, up to 2.84x faster than a fully OpenMP-parallel multiple-precision LU. We also show by measurement that the O(p) advantage of scheme II over scheme I applies only to the GEMM term: modular reduction and CRT grow as O(p^2) and dominate the runtime at the matrix sizes considered here. The implementation is released as open source.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tomonori Kouya. 2026-08-28. Accelerating Multiple-Precision LU Decomposition with Ozaki Scheme II. https://arxiv.org/abs/2609.29554

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Error Estimates for the Arnoldi Approximation of a Matrix Square Root

The Arnoldi process provides an efficient framework for approximating functions of a matrix applied to a vector, i.e., of the form $f(M)\bm{b}$, by repeated matrix-vector multiplications. In this paper, we derive error estimates for approximating the action of a matrix square root using the Arnoldi process, where the integral representation of the error is reformulated in terms of the error for solving the linear system $M\bm{x}=\bm{b}$. The results extend the error analysis of the Lanczos method for Hermitian matrices in [Chen et al., SIAM J. Matrix Anal. Appl., 2022] to non-Hermitian cases and provide an improved bound for the Hermitian case. Furthermore, in practical settings, the matrix may only be available via approximate or structured representations. Motivated by this, we extend the analysis and establish a generalized error bound for perturbed matrices. The numerical results on matrices with different structures demonstrate that our theoretical analysis yields a reliable upper bound. Finally, simulations on large-scale matrices arising in particulate suspensions, represented in hierarchical matrix form, validate the effectiveness and practicality of the approach.

math.NA↗

Study of a TPFA scheme for the stochastic Allen-Cahn problem with constraint through numerical experiments

We study the properties of a fully discrete numerical scheme for the stochastic Allen-Cahn problem with constraint on a bounded polygonal domain in two or three dimensions with homogeneous Neumann boundary condition. The scheme under consideration is of Two Point Flux Approximation (TPFA) type with respect to space and of semi-implicit Euler-Maruyama type with respect to time. The constraint is implemented by a multivalued, subdifferential operator which makes the problem a differential inclusion. This operator is incorporated into the scheme via its Yosida approximation, at the expense of introducing an additional regularization parameter. From previous studies it is known that the time discretization of the Yosida approximation has to be implicit in order to obtain a convergent scheme. Consequently, a non-linear and non-smooth equation has to be solved at each time step and this is challenging for the computation of approximate solutions. In this contribution, we introduce a splitting method to compute the solutions of the TPFA approximations. This method allows us to solve a linear equation in a first step. Then, a piecewise affine projection operator coming from the Yosida approximation can be computed explicitly in a second step. We quantify the error coming from the splitting method and show that the splitting method is accurate. We study properties of the scheme and provide convergence rates through numerical experiments.

math.NA↗

Spectral density estimation for normal matrices

The spectral density estimation problem asks for an algorithm that, given an $n\times n$ matrix $A$, outputs a probability measure that is a good approximation to the uniform distribution on the eigenvalues of $A$, called the spectral density of $A$. This paper considers the setting where $A$ is a large normal matrix that is accessible only through matrix-vector product queries. We provide an algorithm that makes just $m$ matrix-vector queries to $A$ and returns, with high probability, a measure within earth mover's distance $O(1/m+\log m/{\sqrt n})$ of the true spectral density of $A$. We provide a complementary lower bound that any algorithm producing an $\varepsilon$-approximation to the true spectral density for large matrices must make $Ω(1/\varepsilon)$ matrix-vector queries. The lower bound holds even for the more restricted case of real symmetric input matrices. In combination with our upper bound, it shows that spectral density estimation is essentially no harder for complex normal matrices than for real symmetric matrices.

math.NA↗