Search arXivSearch

arXiv · 2511.05035

Modular composition & polynomial GCD in the border of small, shallow circuits

Abstract

Modular composition is the problem of computing the coefficient vector of the polynomial $f(g(x)) \bmod h(x)$, given as input the coefficient vectors of univariate polynomials $f$, $g$, and $h$ over an underlying field $\mathbb{F}$. While this problem is known to be solvable in nearly-linear time over finite fields due to work of Kedlaya & Umans, no such near-linear-time algorithms are known over infinite fields, with the fastest known algorithm being from a recent work of Neiger, Salvy, Schost & Villard that takes $O(n^{1.43})$ field operations on inputs of degree $n$. In this work, we show that for any infinite field $\mathbb{F}$, modular composition is in the border of algebraic circuits with division gates of nearly-linear size and polylogarithmic depth. Moreover, this circuit family can itself be constructed in near-linear time. Our techniques also extend to other algebraic problems, most notably to the problem of computing greatest common divisors of univariate polynomials. We show that over any infinite field $\mathbb{F}$, the GCD of two univariate polynomials can be computed (piecewise) in the border sense by nearly-linear-size and polylogarithmic-depth algebraic circuits with division gates, where the circuits themselves can be constructed in near-linear time. While univariate polynomial GCD is known to be computable in near-linear time by the Knuth--Schönhage algorithm, or by constant-depth algebraic circuits from a recent result of Andrews & Wigderson, obtaining a parallel algorithm that simultaneously achieves polylogarithmic depth and near-linear work remains an open problem of great interest. Our result shows such an upper bound in the setting of border complexity.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert Andrews, Mrinal Kumar, Shanthanu S. Rai. 2026-01-27. Modular composition & polynomial GCD in the border of small, shallow circuits. https://arxiv.org/abs/2511.05035

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bit-counting complexity classes

We define bit-counting complexity classes whose membership depends on the binary profile of the number of accepting paths of non-deterministic polynomial time Turing machines. We study the relationship between this new family of complexity classes and the classical complexity classes. We prove that the complexity class ${\bf PP}$ is contained in our comparison based bit-counting complexity classes ${\bf B_{|0|=|1|}P}$, ${\bf B_{|0|<|1|}P}$ and ${\bf B_{|0|>|1|}P}$. We then show that the comparison based bit-counting complexity classes and the complexity class ${\bf PP}$ are Turing equivalent, that is ${\bf P}^{\bf PP} = {\bf P}^{{\bf B_{|0|=|1|}P}}={\bf P}^{{\bf B_{|0|>|1|}P}}={\bf P}^{{\bf B_{|0|<|1|}P}}$. We then prove that the complexity classes ${\bf NP}$ and ${\bf CoNP}$ are contained in both of our parity based bit-counting complexity classes ${\bf B_{|0| \oplus}P}$ and ${\bf B_{|1| \oplus}P}$. We also show that the Turing closures of the parity based bit-counting complexity classes coincide, that is ${\bf P}^{{\bf B_{|0|\oplus}P}}={\bf P}^{{\bf B_{|1|\oplus}P}}$. We do this by proving that when either parity based bit-counting complexity class is provided as an oracle for a polynomial time Turing machine, then it can simulate the other one, that is ${\bf B_{|1| \oplus}P}\subseteq {\bf P}^{{\bf B_{|0| \oplus}P}}$ and ${\bf B_{|0| \oplus}P}\subseteq {\bf P}^{{\bf B_{|1| \oplus}P}}$.

cs.CC

Formalizing PARITY Circuit Lower Bounds in Lean

We formalize Hastad's PARITY lower bound in Lean using the switching lemma. For every fixed d >= 2, formulas and DAG circuits of computation depth at most d computing PARITY on n inputs require size exp(Omega_d(n^(1/(d-1)))) for all sufficiently large n. This matches the classical upper bound up to constants in the exponent and implies that PARITY is not in nonuniform AC0. We also construct a polynomial-size, logarithmic-depth bounded-fan-in formula family for PARITY, providing a witness to NC1 is not a subset of AC0 for the formalized models. The Lean source code is available at https://github.com/formalcs/circuit-complexity and is checked with Lean 4.33.1 and mathlib 4.33.1.

cs.CC

Constant-Coin Complete-Information Debates for $\mathsf{P}$ with Arbitrarily Small Strong Error

We study complete-information debate systems in which a probabilistic finite-state verifier reads the alternating messages of a prover and a refuter. Demirci, Say, and Yakaryılmaz showed that every language in $\mathsf{P}$ has such debates checkable with a constant number of random bits and arbitrarily small weak error. Their strong-error construction, which also counts nontermination as failure, did not permit arbitrary error reduction. We close this gap: for every $L\in\mathsf{P}$ and every $\varepsilon>0$, there is a constant-space verifier using a constant number of private coin tosses that has perfect completeness and strong error at most $\varepsilon$. The verifier simulates a polynomial-time alternating multihead finite automaton, privately spot-checking one of its input heads. The key observation is that, on a nonmember, the refuter may concede any round in which the prover first misreports a head reading. This ensures termination against every prover when the refuter follows the specified strategy, and permits strong-error reduction by repetition.

cs.CC