Search arXivSearch

arXiv · 2412.06042

Infinite Mixture Models for Improved Modeling of Across-Site Evolutionary Variation

Abstract

Scientific studies in many areas of biology routinely employ evolutionary analyses based on the probabilistic inference of phylogenetic trees from molecular sequence data. Evolutionary processes that act at the molecular level are highly variable, and properly accounting for heterogeneity in evolutionary processes is crucial for more accurate phylogenetic inference. Nucleotide substitution rates and patterns are known to vary among sites in multiple sequence alignments, and such variation can be modeled by partitioning alignments into categories corresponding to different substitution models. Determining $\textit{a priori}$ appropriate partitions can be difficult, however, and better model fit can be achieved through flexible Bayesian infinite mixture models that simultaneously infer the number of partitions, the partition that each site belongs to, and the evolutionary parameters corresponding to each partition. Here, we consider several different types of infinite mixture models, including classic Dirichlet process mixtures, as well as novel approaches for modeling across-site evolutionary variation: hierarchical models for data with a natural group structure, and infinite hidden Markov models that account for spatial patterns in alignments. In analyses of several viral data sets, we find that different types of infinite mixture models emerge as the best choices in different scenarios. To enable these models to scale efficiently to large data sets, we adapt efficient Markov chain Monte Carlo algorithms and exploit opportunities for parallel computing. We implement this infinite mixture modeling framework in BEAST X, a widely-used software package for Bayesian phylogenetic inference.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mandev S. Gill, Guy Baele, Marc A. Suchard, Philippe Lemey. 2024-12-08. Infinite Mixture Models for Improved Modeling of Across-Site Evolutionary Variation. https://arxiv.org/abs/2412.06042

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Mutation Order and Selection Shape Intratumor Heterogeneity in Tumor Evolution

Cancer progression often requires multiple driver mutations, but the same drivers may be acquired in different orders. How these pathways jointly shape tumor clonal structure remains unclear. We develop a multitype branching-process model in which malignant transformation requires two driver mutations, distinguishing malignant cells by mutation order and the independent transformation event that founded their clone. Under a successive exponential approximation, we establish point-process limits for pathway-specific clone sizes and derive a closed-form expression for the limiting expected Simpson's index of the combined malignant population. When both mutation orders yield malignant cells with the same net growth rate, the index decomposes into effective pathway weights, determined by mutation rates and birth-death dynamics at preceding stages, and within-pathway concentration terms, determined by intermediate-to-malignant growth-rate ratios. A driver's effect on heterogeneity thus depends critically on when it is acquired. A strong driver acquired early expands the intermediate lineage and increases the supply of independent malignant founders, whereas the same driver acquired last strengthens the growth and age advantage of early-founded malignant clones. Under additive fitness effects, these opposing mechanisms can produce a non-monotone relationship between selective advantage and clonal concentration. Threshold-like non-additive fitness effects can generate highly concentrated malignant populations, while order-dependent terminal fitness causes the faster-growing pathway to dominate asymptotically. These results show how mutation order, mutational accessibility, selection, and epistasis jointly determine lineage-level intratumor heterogeneity.

q-bio.PE

Phase transitions in microbial lineage trees

Microbial populations exhibit high cell-to-cell variability, which fundamentally shapes population behavior. A striking consequence is the existence of phase transitions, where small genetic or environmental changes trigger abrupt shifts in population dynamics. While biological phase transitions have often been proposed, connecting observed behavior to the underlying physics has remained challenging. We combine population genetics with statistical physics to show how phase transitions arise naturally in microbial populations. We highlight the existence of a first-order transition in a model of bacterial plasmid engineering and find a strict lower bound on the number of plasmids that can be stably maintained in a population.

q-bio.PE

Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk

A well-known phenomenon in statistical analyses of populations of phylogenetic trees in the Billera-Holmes-Vogtmann space is that the topology of the Fréchet mean tree can contain multifurcations (i.e., internal nodes with more than two children), which raises the practical question of whether this reflects a population-level branching structure (hard polytomy) or merely sampling variability in the data (soft polytomy). This is an instance of the more general phenomenon of "stickiness" in non-Euclidean statistics, whereby the sample Fréchet mean in certain non-positively curved stratified spaces becomes permanently trapped in a lower-dimensional stratum. In this work, we identify a particular multidimensional random walk embedded within the Fréchet mean process, and we show that the time at which stickiness occurs is determined by the largest last-passage time above zero of the coordinates of this random walk. Using this representation, we develop a fully nonparametric procedure for estimating the probability that trifurcations in a sample Fréchet mean tree will bifurcate at some future time if more observations are collected. Lastly, we apply our methodology to a problem in phylogenetics where we consider whether an observed trifurcation in the species tree of primates, glires, and tree shrews is genuinely trifurcated at the population level.

q-bio.PE