Search arXivSearch

arXiv · 2606.02408

Structure-Informed Multiple Sequence Alignment: A Formal Model and Hardness Results

Abstract

We formulate a structure-informed multiple sequence alignment problem, denoted MSA-S. The model abstracts biological sequences as strings and structural information as designated position-pairs. It augments a fixed pairwise string score, defined by a fixed non-gap symbol-pair scoring rule and fixed affine gap penalties, with a binary overlap score on designated position-pairs, which can be interpreted as a contact-map overlap score in structural applications. This yields a fixed-score, integer-valued optimization model suitable for complexity-theoretic analysis. Under this formulation, we show that the decision problem MSA-S-DEC is NP-complete for a broad class of fixed pairwise string scoring schemes. We also show that NP-hardness persists even under the restriction that every designated position-pair set is nonempty and the pair-overlap threshold is strictly positive. For the associated scalarized optimization problem MSA-S-OPT(lambda) with any fixed rational constant lambda >= 1, we further show that, under the canonical unit scheme for the non-gap symbol-pair scoring rule, MSA-S-OPT(lambda) admits no polynomial-time approximation scheme (PTAS) even for two input strings (k = 2), unless P = NP. These results establish a formal complexity-theoretic baseline for structure-informed multiple sequence alignment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yoshiki Kanazawa, Naphan Benchasattabuse, Michal Hajdušek, Rodney Van Meter. 2026-06-01. Structure-Informed Multiple Sequence Alignment: A Formal Model and Hardness Results. https://arxiv.org/abs/2606.02408

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Counting Small Induced Subgraphs: Hardness of Symmetry-Based Properties

Jerrum and Meeks (TOCT, JCSS 2015) introduced the counting problems $\text{IndSub}(Φ)$ for fixed graph properties $Φ$: Given an input graph $G$ and $k\in\mathbb N$, count the $k$-vertex subsets $S \subseteq V(G)$ such that the induced subgraph $G[S]$ satisfies $Φ$. For recursively enumerable $Φ$, it is known that $\text{IndSub}(Φ)$ is either #W[1]-hard or fixed-parameter tractable. A direct classification depending on $Φ$ however still remains open. In particular, the status was open for the property of graphs without nontrivial automorphisms, also mentioned in a very recent survey on parameterized counting by Roth (Comput.~Sci.~Rev.~2026). This is a natural property that evades all currently known techniques for proving #W[1]-hardness, including a general toolkit based on Fourier analysis that was very recently introduced by Curticapean and Neuen (SODA~2025). In this paper, we show that counting induced $k$-vertex graphs without nontrivial automorphisms is #W[1]-hard by constructing ``clique scaffolds'', i.e., problem-specific restrictions of the property that enable a reduction from the $k$-clique problem. More generally, we show that for every finite group $Q$, counting $k$-vertex induced subgraphs with automorphism group $Q$ is #W[1]-hard.

cs.CC

Improved Algorithms for the Remote Point Problem

The Remote Point Problem (RPP) is an algorithmic problem that asks, given a linear subspace $L \subseteq \mathbb{F}^n$ of dimension $k$, to deterministically find a vector $v \in \mathbb{F}^n$ far in Hamming distance from $L$. This problem was introduced by Alon, Panigrahy and Yekhanin [APY09], motivated in part by the matrix rigidity approach for proving circuit lower bounds. An algorithm is said to achieve remoteness $d$ if it finds a vector $v$ whose Hamming distance from $L$ is at least $d$. We observe that over the rational numbers, the problem admits a deterministic polynomial-time algorithm that achieves optimal remoteness $n-k$. Over finite fields, we obtain a (modest) improvement of a result of Alon, Panigrahy and Yekhanin [APY09], and give an algorithm that achieves remoteness $Ω\left(\frac{n}{\max\{k, \log n\}} \log n\right)$.

cs.CC

Strong Selective and List-Decoding Direct Product Theorems for Quantum Query Complexity

Quantum strong direct-product theorems for specific functions have been known for nearly two decades. These have been extended to general results for function computation and state generation. The proofs of these results use a version of the multiplicative adversary method that does not naturally extend to relations. Standard strong direct-product theorems apply when algorithms must correctly answer every given question. Prior work extended them to equivalent threshold direct-product theorems, which require answers to all questions but only require that most answers are correct. We focus on two further generalizations. Strong selective direct-products apply to algorithms that adaptively choose, based on what they learn from queries, which questions from a large list to answer. This generalization is relational and useful for proving time-space tradeoffs. We prove a quantum strong selective direct-product theorem for all functions using a new multiplicative adversary formulation for relations that satisfies a strong selective direct product property while being strong enough to capture any query lower bound for functions proven by negative-weights adversaries. This was not previously known even without selectivity. The second generalization is list-decoding direct product problems introduced by Ben-David and Blais for classical randomized query complexity. These allow an algorithm to produce a large list of possible output vectors such that one of them is fully correct. They proved that such theorems hold for classical randomized complexity of all Boolean functions. We prove a quantum analogue of this theorem for all partial Boolean functions. We show that strong list-decoding direct-product theorems are implied by a special case of multiplicative adversaries which we show, via a new reduction, can be obtained from negative-weights adversaries for any Boolean-valued function.

cs.CC