Search arXiv⌕ Search

arXiv · 2610.02931

The Coverage Depth Problem in Distributed DNA Data Storage

Abstract

Random sampling in DNA sequencing produces repeated reads, increasing retrieval latency and sequencing cost. We study the coverage-depth problem for full-message recovery in distributed DNA storage under noiseless uniform sampling, where strands are partitioned among $M$ containers and one strand is independently sampled with replacement from each container per round. For arbitrary linear codes and ordered partitions, we derive exact formulas for the recovery-time distribution and expectation. We prove that MDS codes, whenever they exist, are optimal for every fixed partition, and establish a universal lower bound on the expected total read cost together with its equality conditions. For MDS codes, we identify container-size regimes that yield genuine savings in total reads and regimes that provide only parallelism without changing the asymptotic sequencing cost. For simplex codes, we prove that the $q$-ary simplex code is, up to isomorphism, the unique single-container minimizer among codes with the same parameters, resolving a recent conjecture by Bertuzzo, Ravagnani, and Yaakobi. We further construct a partition attaining the minimum total read cost and derive bounds for intermediate and balanced partitions. These results clarify when distributed sampling reduces latency alone and when it also reduces sequencing cost.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xiangliang Kong, Ohad Elishco, Chen Wang, Tolga M. Duman. 2026-10-02. The Coverage Depth Problem in Distributed DNA Data Storage. https://arxiv.org/abs/2610.02931

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

When isometry and equivalence for skew constacyclic codes coincide

We work in the setting of linear skew constacyclic codes over a commutative base ring $S$. We show that the notions of $(n,σ)$-isometry and $(n,σ)$-equivalence introduced by Ou-azzou et al coincide for most skew $(σ,a)$-constacyclic codes of length $n$, and moreover identify the classes for which the two notions can differ. To prove these results, we first determine all Hamming-weight preserving isomorphisms between their ambient Petit rings which extend some automorphism $τ$ of $S$ that commutes with $σ$. We then prove that when those ambient rings are not associative, these isomorphisms must have degree one, and give examples of non-degree one isomorphisms in the associative case. As a consequence, we provide new definitions of equivalence and isometry that for skew constacyclic codes over finite rings exactly capture all Hamming-weight preserving isomorphisms between their ambient rings, leading to tighter classifications.

cs.IT↗

Coded Computing for Dynamic System via Cartesian Products

This paper studies coded distributed computing (CDC) in a dynamic system in which workers may depart and new clusters may join. The caches of the surviving workers and their existing Reduce assignments stay untouched, while the storage brought by arriving clusters is put to use. In contrast to elastic computing, which targets linear functions, and to dynamic coded caching, which requires placement across all users, the model considered here accommodates general MapReduce tasks under a placement that is partly fixed and partly new. The proposed scheme builds cluster-wise placement delivery arrays (PDAs) and couples them through a Cartesian product. A delivery-aware rule reassigns the abandoned Reduce functions, and a generalized communication PDA allows even workers that hold no Reduce function to serve as coded transmitters. For arbitrary feasible disconnections, the exact load is derived. A file-wise converse for instantly decodable XOR multicasts shows that, in the absence of disconnections, the scheme lies within a factor of two of the best scheme in this multicast class when new clusters arrive, and within a factor of four otherwise, even when the benchmark optimizes over all feasible uncoded placements.

cs.IT↗

Identification and Recovery of Linear Dynamics from Partial Trajectories

We study the joint recovery of an unknown system matrix $A\in\mathbb R^{n\times n}$ and a rank-$r$ initial-state matrix $X_0\in\mathbb R^{n\times m}$ from partial observations of the state matrices $A^tX_0$, for $t=0,\ldots,T$. Each trajectory is observed at a fixed set of state coordinates, which may differ across trajectories. We derive necessary conditions for identifiability, including obstructions arising from incomplete spatial coverage and insufficient dynamical span, as well as lower bounds on the number of observations. We then characterize the Jacobian of the measurement map, account for the intrinsic symmetry of the low-rank factorization, and establish a rank-certificate principle yielding generic local identifiability. When the trajectories whose initial states form a basis of $\operatorname{range}(X_0)$ are fully observed, we obtain necessary and sufficient rank conditions for global recovery, together with explicit reconstruction formulas. We formulate the joint recovery problem as a nonlinear least-squares problem over the system matrix $A$ and the low-rank factors of $X_0$, and solve it using Adam. Numerical experiments on synthetic and real-world data evaluate the recovery performance under different spatial sampling rates and observation horizons.

cs.IT↗