Search arXivSearch

arXiv · 2609.17353

Towards Optimal Prefix-Free Graph Construction: NP-Hardness and Structural Insights

Abstract

Prefix-free parsing provides an efficient way to construct compressed representations of large and repetitive pangenomes and naturally induces a graph representation known as a prefix-free graph. In this work, we initiate a theoretical study of the problem of constructing prefix-free graphs of minimum size, where the size accounts for both the total length of distinct segment labels and the paths representing the input sequences. We show that selecting an optimal set of trigger words is NP-hard, already when triggers consist of single characters. Using a synchronized-code reduction, we extend this hardness result to every fixed trigger length and further show that the problem remains NP-hard over an alphabet of size three. We then establish a structural connection between prefix-free graphs and de Bruijn graphs. In particular, we show that every compacted de Bruijn graph can be realized as a prefix-free graph and derive a hierarchy relating the sizes of minimum pangenomic graphs, minimum prefix-free graphs, compacted de Bruijn graphs, and de Bruijn graphs. Finally, we give an exact fixed-parameter algorithm running in $O(2^q n)$ time, where $q$ is the number of distinct candidate trigger words and $n$ is the total pangenome length. Our results characterize both the computational limitations and the structural properties of optimizing prefix-free graph representations and provide a theoretical foundation for the design of compact graph representations of repetitive pangenomic data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrej Baláž, Alexandru Popa. 2026-09-15. Towards Optimal Prefix-Free Graph Construction: NP-Hardness and Structural Insights. https://arxiv.org/abs/2609.17353

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Counting Small Induced Subgraphs: Hardness of Symmetry-Based Properties

Jerrum and Meeks (TOCT, JCSS 2015) introduced the counting problems $\text{IndSub}(Φ)$ for fixed graph properties $Φ$: Given an input graph $G$ and $k\in\mathbb N$, count the $k$-vertex subsets $S \subseteq V(G)$ such that the induced subgraph $G[S]$ satisfies $Φ$. For recursively enumerable $Φ$, it is known that $\text{IndSub}(Φ)$ is either #W[1]-hard or fixed-parameter tractable. A direct classification depending on $Φ$ however still remains open. In particular, the status was open for the property of graphs without nontrivial automorphisms, also mentioned in a very recent survey on parameterized counting by Roth (Comput.~Sci.~Rev.~2026). This is a natural property that evades all currently known techniques for proving #W[1]-hardness, including a general toolkit based on Fourier analysis that was very recently introduced by Curticapean and Neuen (SODA~2025). In this paper, we show that counting induced $k$-vertex graphs without nontrivial automorphisms is #W[1]-hard by constructing ``clique scaffolds'', i.e., problem-specific restrictions of the property that enable a reduction from the $k$-clique problem. More generally, we show that for every finite group $Q$, counting $k$-vertex induced subgraphs with automorphism group $Q$ is #W[1]-hard.

cs.CC

Improved Algorithms for the Remote Point Problem

The Remote Point Problem (RPP) is an algorithmic problem that asks, given a linear subspace $L \subseteq \mathbb{F}^n$ of dimension $k$, to deterministically find a vector $v \in \mathbb{F}^n$ far in Hamming distance from $L$. This problem was introduced by Alon, Panigrahy and Yekhanin [APY09], motivated in part by the matrix rigidity approach for proving circuit lower bounds. An algorithm is said to achieve remoteness $d$ if it finds a vector $v$ whose Hamming distance from $L$ is at least $d$. We observe that over the rational numbers, the problem admits a deterministic polynomial-time algorithm that achieves optimal remoteness $n-k$. Over finite fields, we obtain a (modest) improvement of a result of Alon, Panigrahy and Yekhanin [APY09], and give an algorithm that achieves remoteness $Ω\left(\frac{n}{\max\{k, \log n\}} \log n\right)$.

cs.CC

Strong Selective and List-Decoding Direct Product Theorems for Quantum Query Complexity

Quantum strong direct-product theorems for specific functions have been known for nearly two decades. These have been extended to general results for function computation and state generation. The proofs of these results use a version of the multiplicative adversary method that does not naturally extend to relations. Standard strong direct-product theorems apply when algorithms must correctly answer every given question. Prior work extended them to equivalent threshold direct-product theorems, which require answers to all questions but only require that most answers are correct. We focus on two further generalizations. Strong selective direct-products apply to algorithms that adaptively choose, based on what they learn from queries, which questions from a large list to answer. This generalization is relational and useful for proving time-space tradeoffs. We prove a quantum strong selective direct-product theorem for all functions using a new multiplicative adversary formulation for relations that satisfies a strong selective direct product property while being strong enough to capture any query lower bound for functions proven by negative-weights adversaries. This was not previously known even without selectivity. The second generalization is list-decoding direct product problems introduced by Ben-David and Blais for classical randomized query complexity. These allow an algorithm to produce a large list of possible output vectors such that one of them is fully correct. They proved that such theorems hold for classical randomized complexity of all Boolean functions. We prove a quantum analogue of this theorem for all partial Boolean functions. We show that strong list-decoding direct-product theorems are implied by a special case of multiplicative adversaries which we show, via a new reduction, can be obtained from negative-weights adversaries for any Boolean-valued function.

cs.CC