Search arXiv⌕ Search

arXiv · 2610.10500

Taxonomic Classification with Complete Tag Arrays

Abstract

Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44\,GB and can be built in minutes on a desktop computer; it classifies a read in 66\,$μ$s with one thread and reaches 93.8\% genus-level accuracy, close to what Cliffy reports for its 9\,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04\,GB, at 75\,$μ$s per read. On the same machine and reads, it is more accurate than Kraken~2 (79.3\%) and Tagger (81.7 to 92.8\%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Travis Gagie, Gonzalo Navarro. 2026-10-07. Taxonomic Classification with Complete Tag Arrays. https://arxiv.org/abs/2610.10500

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Compression with wildcards: All, or all maximum, anticlques of a graph

By definition an anticlique is an independent set of vertices of a graph $G$. By duality all results obtained for anticliques carry over to cliques. (It is for technical reasons that we stick with anticliques throughout.) We display the set $Acl(G)$ of all anticliques of $G$ in a compressed format that uses wildcards. Likewise (albeit less compressed) for the subfamily $MACL(G)\s Acl(G)$ of all maximum-cardinality members. The second task works particularly well for bipartite graphs (in fact for the broader class of König-Egarváry graphs). In this scenario Boolean functions (of type 2-CNF) will be important. Dilworth's lattice of all maximum antichains of a poset also features prominently.

cs.DS↗

Matroid Base Packings: Improved Dynamic Matroid Density and Combinatorics of Tree Packings

Greedy minimum-weight spanning tree packings are an important tool in graph connectivity algorithms. We study the corresponding process of greedy base packing in matroids, following the work of de Vos and Grilnberger. Using a modified version of matroid base packings, we give a fully dynamic $(1 \pm \varepsilon)$-approximation to the matroid density using $O((ρ_{\max}^2\varepsilon^{-2}+ρ_{\max}\varepsilon^{-4})\log^3m_{\max})$ worst-case rank queries per update, where $ρ_{\max}$ upper-bounds the density and $m_{\max}$ upper-bounds the ground set size. Sampling yields a $(1 \pm \varepsilon)$-approximation with high probability against an oblivious adversary using $O(\varepsilon^{-6}\log^6m_{\max})$ worst-case rank queries per update. For graphic matroids, we strengthen the lower bound on the convergence rate of relative edge loads to ideal loads, closing the gap between the lower and upper bounds up to a logarithmic factor. We also show that a packing of $O(λ^5\log m)$ trees contains a tree crossing some minimum cut once, improving the bound $O(λ^7\log^3m)$ of Thorup. In the appendix, we consider a specialization of the greedy base packings to bicircular matroids, which yields a dynamic approximation of the graph density. For this, we develop a dynamic data structure that maintains a minimum-weight maximal pseudoforest.

cs.DS↗

Improved Online Hitting Set Algorithms for Structured and Geometric Set Systems

In the online hitting set problem, sets arrive over time, and the algorithm has to maintain a subset of elements that hit all the sets seen so far. Alon, Awerbuch, Azar, Buchbinder, and Naor (SICOMP 2009) gave an algorithm with competitive ratio $O(\log n \log m)$ for the (general) online hitting set and set cover problems for $m$ sets and $n$ elements; this is known to be tight for efficient online algorithms. Given this barrier for general set systems, we ask: can we break this double-logarithmic phenomenon for online hitting set/set cover on structured and geometric set systems? We provide an $O(\log n \log\log n)$-competitive algorithm for the weighted online hitting set problem on set systems with linear shallow-cell complexity, replacing the double-logarithmic factor in the general result by effectively a single logarithmic term. As a consequence of our results we obtain the first bounds for weighted online hitting set for natural geometric set families, thereby answering open questions regarding the gap between general and geometric weighted online hitting set problems.

cs.DS↗