Search arXivSearch

arXiv · 2307.12712

In-place accumulation of fast multiplication formulae

Abstract

This paper deals with simultaneously fast and in-place algorithms for formulae where the result has to be linearly accumulated: some of the output variables are also input variables, linked by a linear dependency. Fundamental examples include the in-place accumulated multiplication of polynomials or matrices, C+=AB. The difficulty is to combine in-place computations with fast algorithms: those usually come at the expense of (potentially large) extra temporary space, but with accumulation the output variables are not even available to store intermediate values. We first propose a novel automatic design of fast and in-place accumulating algorithms for any bilinear formulae (and thus for polynomial and matrix multiplication) and then extend it to any linear accumulation of a collection of functions. For this, we relax the in-place model to any algorithm allowed to modify its inputs, provided that those are restored to their initial state afterwards. This allows us, in fine, to derive unprecedented in-place accumulating algorithms for fast polynomial multiplications and for Strassen-like matrix multiplications.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jean-Guillaume Dumas, Bruno Grenet. 2024-07-01. In-place accumulation of fast multiplication formulae. https://doi.org/10.1145/3666000.3669671

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parallel Integration over Simple Radical Extensions II: Mixed Towers

In Part I we extended the structure theorems underlying the Risch--Norman (parallel Risch) method to a simple radical extension $L=K(y)$, $y^m=q$, of a differential field $K=F(t_1,\dots,t_n)$ closed under the derivation. Here we remove the closure hypothesis: the radical may occupy any position in the tower, so that the derivatives of the generators above it involve $y$ --- the setting of Bronstein's algorithm for mixed elementary functions. The working ring is $\cA=\cO[t_{j+1},\dots,t_n]$, the integral closure of $F[t_1,\dots,t_n]$ in $L$: a Krull domain, free over the polynomial ring on Trager's basis, so that all factorisation remains in a unique factorisation domain. The denominator of the derivation is no longer an element but a divisor $\fd_D$ on $\cA$, and the valuation lemma takes the unified form $v_P(Dg)=v_P(g)-(1+v_P(\fd_D))$ at normal height-one primes, subsuming the shifts $\{1,e_P\}$ of Part I; the proof localises and requires no cancellation analysis. Stability of the class group and the unit group, $\Cl(\cA)\cong\Cl(\cO)$ and $\cA^*=\cO^*$, splits the admissible logands into $S$-units of $\cO$ --- computed by the machinery of Part I --- and irreducible polynomials moving in the upper variables, whose residues must be constants. We prove degree bounds in the top variable and describe the resulting algorithm, which --- unlike the classical parallel method, whose failure proves nothing --- returns certificates of non-elementarity in two situations: a residue outside the constant field, and, when every bound in force is proved, a residue-free remainder that the linear system shows to be non-exact.

cs.SC

Parallel Integration over Simple Radical Extensions in Mixed Towers: Charlwood's Integrals

We evaluate a SymPy implementation of the parallel Risch-Norman method using Charlwood's 2008 suite of 50 challenging indefinite integrals. The system automatically builds integrand towers, verifies answers by differentiation, and returns correct, verified integrals for 49 problems with zero errors. The single failure, $\int\arcsin(x\sqrt{1-x^2})\,dx$, is proven non-elementary using a holomorphic-remainder certificate, though limited by an unverified completeness hypothesis on a genus-three curve. Part II's degree bounds successfully reduce classical ansatz sizes by two-thirds without impacting running time. The paper details algorithmic mechanisms like $S'$-units, Pell units, and residue computing in tower coordinates. Compared to mature implementations, this untuned SymPy prototype is slower than FriCAS (by a factor of six on the median integral) but more accurate than AXIOM (which returned three wrong answers). Profiling pinpoints performance bottlenecks in nonlinear norm searches and nested number fields, outlining clear targets for optimisation.

cs.SC

Resultant Tools for Parametric Polynomial Systems with Application to Mathematical Biology

To decompose the parameter space of a parametric system of polynomial equations with respect to the number of real solutions that the system attains, first the discriminant variety of the system should be computed. This is an elimination type problem because the discriminant variety is the union of the projections into the parameter space of the intersection of the variety of the original system with one or several new hypersurfaces. Despite the popularity of Gröbner Bases (GBs) for this task, in this work we focus on a lesser used alternative tool, namely resultants. We develop new approaches to build the discriminant variety using different resultant methods such as Dixon resultant and iterated univariate resultants. We discuss how the minimal discriminant variety, the discriminant varieties computed by our methods, and the known border polynomials are related. We demonstrate with numerous examples from population dynamics and chemical reaction network theory that our new approach successfully handles larger size examples than a GB can.

cs.SC