Search arXivSearch

arXiv · 2006.15999

The Number of Repetitions in 2D-Strings

Abstract

The notions of periodicity and repetitions in strings, and hence these of runs and squares, naturally extend to two-dimensional strings. We consider two types of repetitions in 2D-strings: 2D-runs and quartics (quartics are a 2D-version of squares in standard strings). Amir et al. introduced 2D-runs, showed that there are $O(n^3)$ of them in an $n \times n$ 2D-string and presented a simple construction giving a lower bound of $Ω(n^2)$ for their number (TCS 2020). We make a significant step towards closing the gap between these bounds by showing that the number of 2D-runs in an $n \times n$ 2D-string is $O(n^2 \log^2 n)$. In particular, our bound implies that the $O(n^2\log n + \textsf{output})$ run-time of the algorithm of Amir et al. for computing 2D-runs is also $O(n^2 \log^2 n)$. We expect this result to allow for exploiting 2D-runs algorithmically in the area of 2D pattern matching. A quartic is a 2D-string composed of $2 \times 2$ identical blocks (2D-strings) that was introduced by Apostolico and Brimkov (TCS 2000), where by quartics they meant only primitively rooted quartics, i.e. built of a primitive block. Here our notion of quartics is more general and analogous to that of squares in 1D-strings. Apostolico and Brimkov showed that there are $O(n^2 \log^2 n)$ occurrences of primitively rooted quartics in an $n \times n$ 2D-string and that this bound is attainable. Consequently the number of distinct primitively rooted quartics is $O(n^2 \log^2 n)$. Here, we prove that the number of distinct general quartics is also $O(n^2 \log^2 n)$. This extends the rich combinatorial study of the number of distinct squares in a 1D-string, that was initiated by Fraenkel and Simpson (J. Comb. Theory A 1998), to two dimensions. Finally, we show some algorithmic applications of 2D-runs. (Abstract shortened due to arXiv requirements.)

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Panagiotis Charalampopoulos, Jakub Radoszewski, Wojciech Rytter, Tomasz Waleń, Wiktor Zuba. 2020-06-29. The Number of Repetitions in 2D-Strings. https://arxiv.org/abs/2006.15999

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Independent Set Reconfiguration via Dilworth Decompositions

The Token Jumping and Sliding Token problems are fundamental reconfiguration problems defined on the independent sets of an undirected graph. Given two independent sets $I$ and $J$, each of size $k$, these problems ask whether there exists a sequence of elementary operations transforming $I$ into $J$ such that every intermediate configuration is also an independent set of size $k$. Suppose a token is placed on each vertex of $I$: in Sliding Token, an operation moves a token from a vertex $u \in I$ to an adjacent vertex $v \notin I$; in Token Jumping, the token may instead move to any vertex $v \notin I$. While both problems are $\mathsf{PSPACE}$-complete on general graphs, polynomial-time algorithms for one or both variants have been developed for several graph classes, including trees, block graphs, bipartite permutation graphs, cographs, $P_4$-tidy graphs, and interval graphs. In this paper, we prove that both problems are solvable in polynomial time on threshold signed graphs, also known as Dilworth-2 graphs. A graph $G=(V,E)$ is a threshold signed graph if there exist a mapping $a:V\to\mathbb{R}$ and positive real constants $S,T>0$ such that $|a(v)|< \min\{S,T\}$ for all $v \in V$, and for any distinct vertices $u,v\in V$, $\{u,v\}\in E$ if and only if $|a(u)+a(v)|\ge S$ or $|a(u)-a(v)|\ge T$. More generally, we also show that Token Jumping can be solved in time $n^{O(\mathcal{D}(G))}$, where $\mathcal{D}(G)$ denotes the Dilworth number of $G$. Thus, Token Jumping belongs to $\mathsf{XP}$ when parameterised by the Dilworth number. This graph class is a subclass of permutation graphs, for which the complexity of these problems remains open, and is incomparable with the class of bipartite permutation graphs studied by Fox-Epstein et al. (ISAAC, 2015).

cs.DS

Matrix Spencer: Eight Standard Deviations Suffice and an Almost-Linear Time Algorithm for Dense Input

The Matrix Spencer conjecture asserts that for all symmetric matrices $A_1,\ldots,A_n\in\mathbb{R}^{n\times n}$ with $\|A_i\|\le1$ there are signs $\varepsilon_1,\ldots,\varepsilon_n\in\{-1,1\}$ with $\|\sum_{i=1}^n\varepsilon_iA_i\|=O(\sqrt n)$. We prove it: a signing of discrepancy below $8\sqrt n$ always exists. We also give a randomized algorithm that finds a signing of discrepancy below $12\sqrt n$ with failure probability at most $p$. The algorithm uses $n^{3+o(1)}\operatorname{polylog}(1/p)$ arithmetic operations in the real-arithmetic model. This matches the size $n^3$ of the dense input up to subpolynomial factors. In the other direction, we prove that for every $n$ there are collections of symmetric matrices such that every signing has discrepancy at least $(2-o(1))\sqrt{n}$. We present three different proofs of the matrix Spencer conjecture. The key to every proof is a hereditary small-ball estimate. This is a lower bound on the Gaussian measure of the spectral body $\{x\in \mathbb{R}^n:\|\sum_ix_iA_i\|\le R\}$ that holds for every subfamily of the matrices. The other ingredient turns that Gaussian measure into a partial signing. We give three approaches to obtain such a signing. The first one covers the cube by partially signed faces through Gaussian concentration with a constant $7\cdot10^9$. The second proof replaces the covering by a projection lemma with explicit parameters for a constant $156000$. The third proof turns Gaussian measure into signs by a lossless coding, with no union bound. It proves the estimate at the right radius with smooth spectral barriers and certified coefficients. It gives a constant below $7.88$. For algorithms, the main idea is to project Gaussian points onto a smoothed spectral body. The $n^{3+o(1)}$ time algorithm tracks the Gibbs matrix of that body across coordinate-descent steps with sketched increments and random refreshes.

cs.DS

On Deterministically Computing Total Variation Distance via Zonotope Compression

We study deterministic relative approximation of the total variation distance between high-dimensional distributions given by succinct descriptions. We develop an abstract deterministic approximation framework based on representing the total variation distance as a support function of a low-dimensional zonotope. As applications, we obtain FPTASs for several models. Given two mixtures of product distributions over $[q]^n$ with a total of $K$ component distributions, our algorithm approximates their TV-distance within a factor of $1+\varepsilon$ in time $\widetilde O_K(nq(n/\varepsilon)^{2K})$. We also give an FPTAS for mixtures of $n$-step Markov chains over $[q]^n$ with a total of $K$ component distributions, with running time $\widetilde O_K(nq^2(n/\varepsilon)^{2K})$. Finally, for two latent-tree Ising models with the same underlying tree topology, we give an FPTAS for the TV-distance between their leaf marginals in time $O(|V|^{13}\varepsilon^{-12})$.

cs.DS