Search arXiv⌕ Search

arXiv · 2609.32308

SparseDesign: Scaling Exact Coding-Sequence Design

Abstract

Exact optimization of synonymous coding sequences under a joint folding-energy and codon-usage objective is limited by expensive dynamic-programming splits and large working sets. \textsc{SparseDesign} applies candidate sparsification to the multiloop recurrence of a Turner~2004 dangle-0 solver over a weighted codon automaton. A direct branch is retained only when it strictly improves on every partitionable or endpoint-unpaired realization of the same endpoint states. We prove equivalence to the dense recurrence in real arithmetic, under an explicit scalar branch-interface assumption. With $N$ automaton states, edge set $E$ and $Z$ retained candidates, multiloop work is $O(N^2+N|E|+NZ)$; worst-case time remains cubic for bounded-width automata and total memory remains quadratic. Endpoint ownership permits parallel candidate construction without locks. While synthetic stress families can benefit little from sparsification and exhibit near-quadratic candidate growth, natural proteins show substantial candidate-count reductions. In our 7,600-task campaign, the 2,000-protein human-table panel has median retention of only 3.53\% at $λ=0$ and 2.15\% at $λ=4$, corresponding to approximately 28.3-fold and 46.4-fold reductions relative to all feasible direct intervals. The primary performance experiments use an AMD EPYC 7313 server. For human Dp427c (11,031 nt, $λ=0$), 16-thread packed \textsc{SparseDesign} achieves five-run medians of 236.54 seconds wall-clock time and 14.43 GiB peak RSS. Compared with the single-thread local dense LinearDesign fork on the same server (4,912 seconds, 402.10 GiB RSS), this gives a 20.8-fold wall-clock speedup and a 27.9-fold peak-memory reduction. On a Core i9-14900KF commodity PC with 64 GiB RAM, the same input, layout and thread count achieve 126.42 seconds and 14.43 GiB RSS.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hao Lin, Jingjin Yu. 2026-09-26. SparseDesign: Scaling Exact Coding-Sequence Design. https://arxiv.org/abs/2609.32308

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Capacitated Partition Vertex Cover and Partition Edge Cover

Our first focus is the Capacitated Partition Vertex Cover (C-PVC) problem in hypergraphs. In C-PVC, we are given a hypergraph with capacities on its vertices and a partition of the hyperedge set into $ω$ distinct groups. The objective is to select a minimum size subset of vertices that satisfies two main conditions: (1) in each group, the total number of covered hyperedges meets a specified threshold, and (2) the number of hyperedges assigned to any vertex respects its capacity constraint. A covered hyperedge is required to be assigned to a selected vertex that belongs to the hyperedge. This formulation generalizes classical Vertex Cover, Partial Vertex Cover, and Partition Vertex Cover. We investigate two variants: soft capacitated (multiple copies of a vertex are allowed) and hard capacitated (each vertex can be chosen at most once). Let $f$ denote the rank of the hypergraph (i.e., the maximum number of vertices contained in any single hyperedge). Our main contributions are: $(i)$ an $(f+1)$-approximation algorithm for the weighted soft-capacitated C-PVC problem, which runs in polynomial time for constant \(ω\), and $(ii)$ an $(f+ε)$-approximation algorithm for the unweighted hard-capacitated C-PVC problem, which runs in $n^{O(ω/ε)}$ time. We also study a natural generalization of the edge cover problem, the \emph{Weighted Partition Edge Cover} (W-PEC) problem, where each edge has an associated weight, and the vertex set is partitioned into groups. For each group, the goal is to cover at least a specified number of vertices using incident edges, while minimizing the total weight of the selected edges. We present the first exact polynomial-time algorithm for the weighted case, improving runtime from $O(ωn^3)$ to $O(mn+n^2 \log n)$ and simplifying the algorithmic structure over prior unweighted approaches.

cs.DS↗

Randomization for Faster Exact Optimization of Discounted Markov Decision Processes

We provide faster deterministic and randomized algorithms for exactly solving discounted Markov Decision Processes (DMDPs). We obtain our results by efficiently reducing computing optimal values and policies in DMDPs to the easier tasks of policy evaluation and computing approximately optimal values in DMDPs. We provide both a straightforward deterministic reduction and a more efficient randomized variant that, together with advances in approximately solving DMDPs, yield our results.

cs.DS↗

An ETH-Tight, Constructive FPT Algorithm for the Cone and Polytope Intersection Problem

In a landmark paper, Goemans and Rothvoss (2020) established an XP algorithm running in time $\text{enc}(P)^{2^{O(d)}} \cdot \text{enc}(Q)^{O(1)}$ for the Cone and Polytope Intersection problem: finding a vector $y \in \text{int.cone}(P \cap \mathbb{Z}^d) \cap Q$ together with a sparse certificate $λ\in \mathbb{Z}_{\ge 0}^{P \cap \mathbb{Z}^d}$ supported on at most $2^{2d+1}$ generators, where $P \subseteq \mathbb{R}^d$ is a bounded rational polyhedron and $Q \subseteq \mathbb{R}^d$ is an arbitrary rational polyhedron. For high-multiplicity bin packing, this gives a running time of ${|I|}^{2^{O(d)}}$, where $|I|$ denotes the encoding length of the input. Recently, Koana and Kumabe (2026) proved that the decision variant of this problem is fixed-parameter tractable (FPT) parameterized by the number of item types $d$ with running time $2^{d^{O(d)}} \cdot {|I|}^{O(1)} = 2^{2^{O(d \log d)}} \cdot {|I|}^{O(1)}$. In this work, we generalize the framework of Koana and Kumabe from standard bin packing to the full Cone and Polytope Intersection Problem of Goemans and Rothvoss, directly encompassing high-multiplicity bin packing, point-in-cone, and scheduling. Secondly, by combining Carathéodory-type integer cone bounds (Eisenbrand and Shmonin, 2006) with active support enumeration, we reduce the running time to: $$2^{2^{O(d)}} \cdot (\text{enc}(P) + \text{enc}(Q))^{O(1)}.$$ Under the Exponential Time Hypothesis (ETH), the double-exponential lower bound of Kowalik, Lassota, Majewski, Pilipczuk, and Sokołowski (2024) for point-in-cone and Jansen, Ohnesorge, and Pirotton (2026) for high-multiplicity bin packing implies that this parameter dependence is asymptotically optimal. Finally, we provide an explicit decompression algorithm that extracts a solution with sparse support $|\text{supp}(λ)| \le 2^{2d+1}$ in single-exponential FPT time.

cs.DS↗