Search arXiv⌕ Search

arXiv · 2610.05696

Substring Edit Correcting Codes and Optimal Single Burst-Deletion Correcting Codes

Abstract

A $k$-substring edit in a sequence first deletes a substring of length at most $k$, and then inserts a sequence of length at most $k$ at the same position. A code that can correct a $k$-substring edit is called a $k$-substring edit code. In this paper, we develop a new localization method and use it to construct a $q$-ary $k$-substring edit correcting code with $\log n+8\log\log n+o(\log\log n)$ bits of redundancy for any fixed $q\ge2$ and $k\ge1$, where $n$ is the code length. For the binary alphabet, this improves upon the redundancy $\log n+16k\log\log n+o(\log\log n)$ obtained by Li \emph{et al}. When the deleted substring and the inserted sequence have different lengths, we further construct codes with redundancy $\log n+O_{q,k}(1)$, which is optimal up to an additive constant. As corollaries, for all fixed $q\ge2$ and $1\le t\le T$, we obtain $q$-ary $(\le t)$-burst-deletion correcting codes and $(t,T)$-localized deletion correcting codes with redundancies $\log n+O_{q,t}(1)$ and $\log n+O_{q,T}(1)$, respectively. To the best of our knowledge, these are the first constructions attaining optimal redundancy up to an additive constant for these two deletion models over the full range of fixed parameters. For $(\le t)$-burst-deletion correction with $t\ge2$, such redundancy had previously been achieved only for $q=t=2$ by Levenshtein in 1967.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zuo Ye, Gennian Ge. 2026-10-05. Substring Edit Correcting Codes and Optimal Single Burst-Deletion Correcting Codes. https://arxiv.org/abs/2610.05696

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parametric and structure-aware information theory for multi-scale analysis of composition

Compositional data is common across the natural and social sciences, requiring methods to measure the diversity of compositions and the heterogeneity of collections of them. Via Bregman geometry, a strictly concave diversity index induces a measure of heterogeneity that decomposes across scales. A canonical example is Shannon entropy inducing mutual information. We develop this framework by incorporating category similarity into $α$-logarithmic entropy, a parametric diversity index. We disprove an existing general concavity result and prove that, for $α=3$, the region of strict concavity is always convex, complementing previous results for $α=1$ and $2$. The resulting geometry provides methods to compare compositions substantially faster than optimal transport. Applied to occupation compositions across England and Wales, they reveal distinct regionalisations based on different notions of occupation relatedness. Applied to ecological data, they recover established patterns of functional and taxonomic $β$-diversity, and reveal sensitivity to the diversity parameter.

cs.IT↗

Improved Characterization of the Memory-Rate Tradeoff for Demand-Private Coded Caching With Multiple Demands

We consider a coded caching problem with multiple demands under a privacy constraint. In this problem, a server with access to $N$ files serves $K$ users over a shared link, and each user requests $L$ distinct files. The privacy constraint requires that each user obtain no information about the demands of the other users. We first propose a new achievable scheme for arbitrary $N$, $K$, and $L$ by applying a novel transformation to a suitable non-private coded caching scheme. We then derive a new converse bound, and show that the proposed scheme is order optimal within a multiplicative factor of $6$ of this bound. Finally, for the case of two users, we completely characterize the optimal memory-rate tradeoff for arbitrary $N$ and $L$ through three new achievable schemes and matching converse bounds.

cs.IT↗

Learning LDPC codes with density evolution over relaxed protographs

We consider the design of low-density parity-check (LDPC) codes for a given iterative decoder. While LDPC performance can be evaluated using simulation, density evolution (DE), or EXIT-chart analysis, selecting a parity-check matrix (PCM) remains a difficult combinatorial optimization problem. Existing approaches often rely on population-based search, random mutations, or genetic algorithms, which require careful tuning and incur high computational cost. Recent gradient descent (GD)-based methods optimize relaxed PCMs by differentiating through decoder simulations, but rely on noisy Monte Carlo estimates, line searches over soft matrix representations, and remain costly for long codes. Moreover, the loss is typically evaluated only at integer-valued PCMs. We focus on long protograph-based LDPC codes and propose a deterministic GD-based framework operating directly on a relaxed protograph representation. The loss is based on DE bit error rate (BER) and can be evaluated directly for relaxed protographs. To justify this relaxation, we associate the relaxed representation with an ensemble of binary PCMs and show that the proposed relaxed DE yields the ensemble-averaged DE performance. The resulting procedure supports standard GD optimization and achieves fast, reliable convergence through deterministic DE evaluation and informative gradients. Numerical results show that the optimized protographs outperform 5G LDPC codes with matching dimensions.

cs.IT↗