Search arXivSearch

arXiv · 2302.13600

Fast in-place accumulation

Abstract

This paper deals with simultaneously fast and in-place algorithms for formulae where the result has to be linearly accumulated: some output variables are also input variables, linked by a linear dependency. Fundamental examples include the in-place accumulated multiplication of polynomials or matrices (that is with only $O(1)$ extra space). The difficulty is to combine in-place computations with fast algorithms: those usually come at the expense of (potentially large) extra temporary space. We first propose a novel automatic design of fast and in-place accumulating algorithms for any bilinear formulae (and thus for polynomial and matrix multiplication) and then extend it to any linear accumulation of a collection of functions. For this, we relax the in-place model to any algorithm allowed to modify its inputs, provided that those are restored to their initial state afterwards. This allows us to derive in-place accumulating algorithms for fast polynomial multiplications and for Strassen-like matrix multiplications. We then consider the simultaneously fast and in-place computation of the Euclidean polynomial remainder $R = A \bmod B$. If $A$ and $B$ have respective degree $m+n$ and $n$, and $M(k)$ denotes the complexity of a (not-in-place) algorithm to multiply two degree-$k$ polynomials, our algorithm uses at most $O((n/m) M(m)\log(m))$ arithmetic operations. If $M(n) = Θ(n^{1+ε})$ for some $ε>0$, then our algorithms do match the not-in-place complexity bound of $O((n/m) M(m))$. We also propose variants that compute - still in-place and with the same complexity bounds - $A = A \bmod B$, $R += A \bmod B$ and $R += AC \bmod B$, that is multiplication in a finite field extension. To achieve this, we develop techniques for Toeplitz matrix operations, for generalized convolutions, short product and power series division and remainder whose output is also part of the input.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jean-Guillaume Dumas, Bruno Grenet. 2025-11-06. Fast in-place accumulation. https://doi.org/10.1016/j.jsc.2025.102523

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parallel Integration over Simple Radical Extensions II: Mixed Towers

In Part I we extended the structure theorems underlying the Risch--Norman (parallel Risch) method to a simple radical extension $L=K(y)$, $y^m=q$, of a differential field $K=F(t_1,\dots,t_n)$ closed under the derivation. Here we remove the closure hypothesis: the radical may occupy any position in the tower, so that the derivatives of the generators above it involve $y$ --- the setting of Bronstein's algorithm for mixed elementary functions. The working ring is $\cA=\cO[t_{j+1},\dots,t_n]$, the integral closure of $F[t_1,\dots,t_n]$ in $L$: a Krull domain, free over the polynomial ring on Trager's basis, so that all factorisation remains in a unique factorisation domain. The denominator of the derivation is no longer an element but a divisor $\fd_D$ on $\cA$, and the valuation lemma takes the unified form $v_P(Dg)=v_P(g)-(1+v_P(\fd_D))$ at normal height-one primes, subsuming the shifts $\{1,e_P\}$ of Part I; the proof localises and requires no cancellation analysis. Stability of the class group and the unit group, $\Cl(\cA)\cong\Cl(\cO)$ and $\cA^*=\cO^*$, splits the admissible logands into $S$-units of $\cO$ --- computed by the machinery of Part I --- and irreducible polynomials moving in the upper variables, whose residues must be constants. We prove degree bounds in the top variable and describe the resulting algorithm, which --- unlike the classical parallel method, whose failure proves nothing --- returns certificates of non-elementarity in two situations: a residue outside the constant field, and, when every bound in force is proved, a residue-free remainder that the linear system shows to be non-exact.

cs.SC

Parallel Integration over Simple Radical Extensions in Mixed Towers: Charlwood's Integrals

We evaluate a SymPy implementation of the parallel Risch-Norman method using Charlwood's 2008 suite of 50 challenging indefinite integrals. The system automatically builds integrand towers, verifies answers by differentiation, and returns correct, verified integrals for 49 problems with zero errors. The single failure, $\int\arcsin(x\sqrt{1-x^2})\,dx$, is proven non-elementary using a holomorphic-remainder certificate, though limited by an unverified completeness hypothesis on a genus-three curve. Part II's degree bounds successfully reduce classical ansatz sizes by two-thirds without impacting running time. The paper details algorithmic mechanisms like $S'$-units, Pell units, and residue computing in tower coordinates. Compared to mature implementations, this untuned SymPy prototype is slower than FriCAS (by a factor of six on the median integral) but more accurate than AXIOM (which returned three wrong answers). Profiling pinpoints performance bottlenecks in nonlinear norm searches and nested number fields, outlining clear targets for optimisation.

cs.SC

Resultant Tools for Parametric Polynomial Systems with Application to Mathematical Biology

To decompose the parameter space of a parametric system of polynomial equations with respect to the number of real solutions that the system attains, first the discriminant variety of the system should be computed. This is an elimination type problem because the discriminant variety is the union of the projections into the parameter space of the intersection of the variety of the original system with one or several new hypersurfaces. Despite the popularity of Gröbner Bases (GBs) for this task, in this work we focus on a lesser used alternative tool, namely resultants. We develop new approaches to build the discriminant variety using different resultant methods such as Dixon resultant and iterated univariate resultants. We discuss how the minimal discriminant variety, the discriminant varieties computed by our methods, and the known border polynomials are related. We demonstrate with numerous examples from population dynamics and chemical reaction network theory that our new approach successfully handles larger size examples than a GB can.

cs.SC