Search arXivSearch

arXiv · 2108.08143

Effective and scalable clustering of SARS-CoV-2 sequences

Abstract

SARS-CoV-2, like any other virus, continues to mutate as it spreads, according to an evolutionary process. Unlike any other virus, the number of currently available sequences of SARS-CoV-2 in public databases such as GISAID is already several million. This amount of data has the potential to uncover the evolutionary dynamics of a virus like never before. However, a million is already several orders of magnitude beyond what can be processed by the traditional methods designed to reconstruct a virus's evolutionary history, such as those that build a phylogenetic tree. Hence, new and scalable methods will need to be devised in order to make use of the ever increasing number of viral sequences being collected. Since identifying variants is an important part of understanding the evolution of a virus, in this paper, we propose an approach based on clustering sequences to identify the current major SARS-CoV-2 variants. Using a $k$-mer based feature vector generation and efficient feature selection methods, our approach is effective in identifying variants, as well as being efficient and scalable to millions of sequences. Such a clustering method allows us to show the relative proportion of each variant over time, giving the rate of spread of each variant in different locations -- something which is important for vaccine development and distribution. We also compute the importance of each amino acid position of the spike protein in identifying a given variant in terms of information gain. Positions of high variant-specific importance tend to agree with those reported by the USA's Centers for Disease Control and Prevention (CDC), further demonstrating our approach.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sarwan Ali, Tamkanat-E-Ali, Muhammad Asad Khan, Imdadullah Khan, Murray Patterson. 2021-10-12. Effective and scalable clustering of SARS-CoV-2 sequences. https://arxiv.org/abs/2108.08143

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Phase transitions in microbial lineage trees

Microbial populations exhibit high cell-to-cell variability, which fundamentally shapes population behavior. A striking consequence is the existence of phase transitions, where small genetic or environmental changes trigger abrupt shifts in population dynamics. While biological phase transitions have often been proposed, connecting observed behavior to the underlying physics has remained challenging. We combine population genetics with statistical physics to show how phase transitions arise naturally in microbial populations. We highlight the existence of a first-order transition in a model of bacterial plasmid engineering and find a strict lower bound on the number of plasmids that can be stably maintained in a population.

q-bio.PE

Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk

A well-known phenomenon in statistical analyses of populations of phylogenetic trees in the Billera-Holmes-Vogtmann space is that the topology of the Fréchet mean tree can contain multifurcations (i.e., internal nodes with more than two children), which raises the practical question of whether this reflects a population-level branching structure (hard polytomy) or merely sampling variability in the data (soft polytomy). This is an instance of the more general phenomenon of "stickiness" in non-Euclidean statistics, whereby the sample Fréchet mean in certain non-positively curved stratified spaces becomes permanently trapped in a lower-dimensional stratum. In this work, we identify a particular multidimensional random walk embedded within the Fréchet mean process, and we show that the time at which stickiness occurs is determined by the largest last-passage time above zero of the coordinates of this random walk. Using this representation, we develop a fully nonparametric procedure for estimating the probability that trifurcations in a sample Fréchet mean tree will bifurcate at some future time if more observations are collected. Lastly, we apply our methodology to a problem in phylogenetics where we consider whether an observed trifurcation in the species tree of primates, glires, and tree shrews is genuinely trifurcated at the population level.

q-bio.PE

Coexistence coalitions in propagule disperser quasi-communities

Many natural ecosystems harbor large numbers of coexisting species competing for far fewer distinct resources, in apparent defiance of the competitive exclusion principle. Various mechanisms have been proposed to explain this apparent paradox, often pertaining to organisms with a two-stage sessile--propagule life cycle. Here we develop a stochastic model class for such propagule disperser communities that combines competition--colonization trade-offs, spatial heterogeneity, demographic stochasticity, as well as inherited trait variation, and recover several classical models as special or limiting cases. Using bifurcation analysis, we classify equilibrium coalitions near the extinction threshold and give sufficient conditions for their realization by macroscopic equilibria away from the threshold, bypassing the costly numerical computation of the actual equilibrium states. Illustrative examples examine the resulting trait distributions and coalition patterns, demonstrating the interactive effects of different coexistence mechanisms.

q-bio.PE