Search arXiv⌕ Search

arXiv · 2508.03930

Counting Distinct Square Substrings in Sublinear Time

Abstract

We show that the number of distinct squares in a packed string of length $n$ over an alphabet of size $σ$ can be computed in $O(n/\log_σn)$ time in the word-RAM model. This paper is the first to introduce a sublinear-time algorithm for counting squares in the packed setting. The packed representation of a string of length $n$ over an alphabet of size $σ$ is given as a sequence of $O(n/\log_σn)$ machine words in the word-RAM model (a machine word consists of $ω\ge \log_2 n$ bits). Previously, it was known how to count distinct squares in $O(n)$ time [Gusfield and Stoye, JCSS 2004], even for a string over an integer alphabet [Crochemore et al., TCS 2014; Bannai et al., CPM 2017; Charalampopoulos et al., SPIRE 2020]. We use the techniques for extracting squares from runs described by Crochemore et al. [TCS 2014]. However, the packed model requires novel approaches. We need an $O(n/\log_σn)$-sized representation of all long-period runs (runs with period $Ω(\log_σn)$) which allows for a sublinear-time counting of the -- potentially linearly-many -- implied squares. The long-period runs with a string period that is periodic itself (called layer runs) are an obstacle, since their number can be $Ω(n)$. The number of all other long-period runs is $O(n/\log_σn)$ and we can construct an implicit representation of all long-period runs in $O(n/\log_σn)$ time by leveraging the insights of Amir et al. [ESA 2019]. We count squares in layer runs by exploiting combinatorial properties of pyramidally-shaped groups of layer runs. Another difficulty lies in computing the locations of Lyndon roots of runs in packed strings, which is needed for grouping runs that may generate equal squares. To overcome this difficulty, we introduce sparse-Lyndon roots which are based on string synchronizers [Kempa and Kociumaka, STOC 2019].

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Panagiotis Charalampopoulos, Manal Mohamed, Jakub Radoszewski, Wojciech Rytter, Tomasz Waleń, Wiktor Zuba. 2025-09-02. Counting Distinct Square Substrings in Sublinear Time. https://arxiv.org/abs/2508.03930

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Path Enumeration by Position-Visit Counts in Recombining Trinomial Trees

Recombining trinomial trees are a workhorse for modeling discrete-event systems in option pricing, logistics, and feedback control. Because each node stores a state-dependent quantity, a depth-$D$ tree contains $3^D$ raw trajectories, making exhaustive enumeration rapidly infeasible. However, when each node's value depends only on its position, a raw trajectory's aggregate is determined by its position-visit counts. We call these count vectors cardinality tuples and decompose the admissible tuples into weak-composition mass layers. Leveraging these structures, we introduce a mass-shifting enumeration algorithm that slides integer ``masses'' through cardinality tuples to generate exactly one representative of each path-equivalence class, while the accompanying weak-composition bijections yield exact counting formulas for the generated families. This suppresses redundant raw-path orderings a priori rather than enumerating and deduplicating them afterward. For the full-tuple implementation, we prove an output-sensitive running-time bound at each fixed endpoint, together with a uniform worst-case upper bound $\mathscr{O}(D2^D)$ and an exact worst-case exponential growth base of $2$, compared with base $3$ for exhaustive raw-path enumeration. Thus the construction achieves a provable exponential reduction in the enumeration space, up to polynomial factors. The same framework also recovers the information compressed by the equivalence classes: we derive an exact degeneracy formula for the number of raw paths represented by every cardinality tuple. We further prove that the nonnegative return specialization is exactly the classical Motzkin family, recover its recursive and generating-function structure and the Dyck specialization, and derive a multivariate occupation-profile $J$-fraction whose coefficients recover the corresponding cardinality-tuple degeneracies.

cs.DS↗

The Longest Common Bitonic Subsequence: Match-Sensitive Algorithms and Conditional Hardness

The longest common bitonic subsequence problem asks for a longest common subsequence of two ordered sequences whose values strictly increase and then strictly decrease; either phase may be empty. We formulate the problem through increasing and decreasing endpoint values at matching position pairs. This gives a constructive quadratic baseline and a matchsensitive algorithm based on two standard dominance-maximum passes. Its time is the sum of an input-sorting term and the number of matches times a squared logarithmic factor. We state the endpoint interface that permits reuse of increasing subsequence algorithms, and distinguish this specialization from new range searching machinery. A linear-size padding reduction transfers the conditional strongly subquadratic lower bound for longest common increasing subsequence to the bitonic problem. Reproducible implementations, exhaustive small-instance checks, and newly measured synthetic experiments document correctness and the practical tradeoff between sparse and dense processing.

cs.DS↗

Locally Approximating the Top Eigenvector of Bounded Entry Matrices

We provide a local computation algorithm to approximate the top eigenvector $x \in \mathbb{R}^n$ of a symmetric matrix $A \in \mathbb{R}^{n \times n}$ with entries between $-1$ and $1$, building on the work of Swartworth and Woodruff [SODA 25] who show how to approximate the eigenvalues up to additive-$\varepsilon n$ error using $\tilde{O}(1/\varepsilon^4)$ queries. Our local computation algorithm has a preprocessing complexity of $\tilde{O}(1/\varepsilon^4)$ and per-coordinate query complexity of $\tilde{O}(1/\varepsilon^2)$ for an additive-$\varepsilon n$ approximation whenever {$|λ_{\min}(A)| = O(λ_{\max}(A))$. When $λ_{\min}(A)$ greatly exceeds $λ_{\max}(A)$, our complexity degrades to at most $\tilde{O}(1/\varepsilon^{6.\overline{6}})$ in preprocessing and $\tilde{O}(1/\varepsilon^{3.\overline{3}})$ per query. Furthermore, we show a lower bound of $Ω(n/\varepsilon^2)$ on the total number of queries needed to output an approximately top eigenvector (implying that the per-coordinate query complexity of $Ω(1/\varepsilon^2)$ is necessary). As an application, we use our algorithm to provide local computation algorithms for the sparsest-cut and max-cut problems in the dense graph model of Goldreich, Goldwasser, Ron [JACM 98]. By accessing the top eigenvectors (of an approximate normalized adjacency), we implement local versions of Cheeger's inequality and Trevisan's algorithm [SICOMP 12] to obtain "square-root-opt" approximations in polynomial time (as opposed to exponential-in-$\text{poly}(1/\varepsilon)$ time which is incurred in Goldreich, Goldwasser, Ron.

cs.DS↗