Search arXiv⌕ Search

arXiv · 2609.31501

Fundamental Limits of Sequence Reconstruction Problems in Immunogenomics

Abstract

The goal of personalized immunogenomics is to recover an individual's germline immunoglobulin gene segments from 'repertoire sequences' altered by trimming, extension, and mutation. In this work, we study the fundamental trace complexity (or the number of samples/traces required for accurate reconstruction) of D-gene reconstruction under three biologically motivated trace-generation models introduced by Bhardwaj et al. (2021) and develop practical algorithms for reconstruction problems left open in the original work. First, for the TrimSuffixAndExtend model, we establish the optimal trace complexity to be {Θ(n)}, and develop a low-complexity Prefix-Filtered Mode (PFM) decoder that achieves this scaling. Second, for the closely related two-sided TrimAndExtend model, we show that the optimal trace complexity is instead {Θ(n^2)}, and is achieved by the simple Bit-Wise Mode (BWM) decoder. Third, for the SuffixExtend-t(TrimSuffix) model, we establish polynomially separated lower and upper bounds on trace complexity. Our results follow from information-theoretic lower bounds coupled with tight analyses of the proposed reconstruction algorithms.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jaswanthi Mandalapu, Deepak Charan, V. Arvind Rameshwar, Nir Weinberger. 2026-09-25. Fundamental Limits of Sequence Reconstruction Problems in Immunogenomics. https://arxiv.org/abs/2609.31501

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parametric and structure-aware information theory for multi-scale analysis of composition

Compositional data is common across the natural and social sciences, requiring methods to measure the diversity of compositions and the heterogeneity of collections of them. Via Bregman geometry, a strictly concave diversity index induces a measure of heterogeneity that decomposes across scales. A canonical example is Shannon entropy inducing mutual information. We develop this framework by incorporating category similarity into $α$-logarithmic entropy, a parametric diversity index. We disprove an existing general concavity result and prove that, for $α=3$, the region of strict concavity is always convex, complementing previous results for $α=1$ and $2$. The resulting geometry provides methods to compare compositions substantially faster than optimal transport. Applied to occupation compositions across England and Wales, they reveal distinct regionalisations based on different notions of occupation relatedness. Applied to ecological data, they recover established patterns of functional and taxonomic $β$-diversity, and reveal sensitivity to the diversity parameter.

cs.IT↗

Improved Characterization of the Memory-Rate Tradeoff for Demand-Private Coded Caching With Multiple Demands

We consider a coded caching problem with multiple demands under a privacy constraint. In this problem, a server with access to $N$ files serves $K$ users over a shared link, and each user requests $L$ distinct files. The privacy constraint requires that each user obtain no information about the demands of the other users. We first propose a new achievable scheme for arbitrary $N$, $K$, and $L$ by applying a novel transformation to a suitable non-private coded caching scheme. We then derive a new converse bound, and show that the proposed scheme is order optimal within a multiplicative factor of $6$ of this bound. Finally, for the case of two users, we completely characterize the optimal memory-rate tradeoff for arbitrary $N$ and $L$ through three new achievable schemes and matching converse bounds.

cs.IT↗

Learning LDPC codes with density evolution over relaxed protographs

We consider the design of low-density parity-check (LDPC) codes for a given iterative decoder. While LDPC performance can be evaluated using simulation, density evolution (DE), or EXIT-chart analysis, selecting a parity-check matrix (PCM) remains a difficult combinatorial optimization problem. Existing approaches often rely on population-based search, random mutations, or genetic algorithms, which require careful tuning and incur high computational cost. Recent gradient descent (GD)-based methods optimize relaxed PCMs by differentiating through decoder simulations, but rely on noisy Monte Carlo estimates, line searches over soft matrix representations, and remain costly for long codes. Moreover, the loss is typically evaluated only at integer-valued PCMs. We focus on long protograph-based LDPC codes and propose a deterministic GD-based framework operating directly on a relaxed protograph representation. The loss is based on DE bit error rate (BER) and can be evaluated directly for relaxed protographs. To justify this relaxation, we associate the relaxed representation with an ensemble of binary PCMs and show that the proposed relaxed DE yields the ensemble-averaged DE performance. The resulting procedure supports standard GD optimization and achieves fast, reliable convergence through deterministic DE evaluation and informative gradients. Numerical results show that the optimized protographs outperform 5G LDPC codes with matching dimensions.

cs.IT↗