Search arXivSearch

arXiv · 1004.5465

Selective Constraints on Amino Acids Estimated by a Mechanistic Codon Substitution Model with Multiple Nucleotide Changes

Abstract

Empirical substitution matrices represent the average tendencies of substitutions over various protein families by sacrificing gene-level resolution. We develop a codon-based model, in which mutational tendencies of codon, a genetic code, and the strength of selective constraints against amino acid replacements can be tailored to a given gene. First, selective constraints averaged over proteins are estimated by maximizing the likelihood of each 1-PAM matrix of empirical amino acid (JTT, WAG, and LG) and codon (KHG) substitution matrices. Then, selective constraints specific to given proteins are approximated as a linear function of those estimated from the empirical substitution matrices. Akaike information criterion (AIC) values indicate that a model allowing multiple nucleotide changes fits the empirical substitution matrices significantly better. Also, the ML estimates of transition-transversion bias obtained from these empirical matrices are not so large as previously estimated. The selective constraints are characteristic of proteins rather than species. However, their relative strengths among amino acid pairs can be approximated not to depend very much on protein families but amino acid pairs, because the present model, in which selective constraints are approximated to be a linear function of those estimated from the JTT/WAG/LG/KHG matrices, can provide a good fit to other empirical substitution matrices including cpREV for chloroplast proteins and mtREV for vertebrate mitochondrial proteins. The present codon-based model with the ML estimates of selective constraints and with adjustable mutation rates of nucleotide would be useful as a simple substitution model in ML and Bayesian inferences of molecular phylogenetic trees, and enables us to obtain biologically meaningful information at both nucleotide and amino acid levels from codon and protein sequences.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sanzo Miyazawa. 2011-08-30. Selective Constraints on Amino Acids Estimated by a Mechanistic Codon Substitution Model with Multiple Nucleotide Changes. https://doi.org/10.1371/journal.pone.0017244

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Mutation Order and Selection Shape Intratumor Heterogeneity in Tumor Evolution

Cancer progression often requires multiple driver mutations, but the same drivers may be acquired in different orders. How these pathways jointly shape tumor clonal structure remains unclear. We develop a multitype branching-process model in which malignant transformation requires two driver mutations, distinguishing malignant cells by mutation order and the independent transformation event that founded their clone. Under a successive exponential approximation, we establish point-process limits for pathway-specific clone sizes and derive a closed-form expression for the limiting expected Simpson's index of the combined malignant population. When both mutation orders yield malignant cells with the same net growth rate, the index decomposes into effective pathway weights, determined by mutation rates and birth-death dynamics at preceding stages, and within-pathway concentration terms, determined by intermediate-to-malignant growth-rate ratios. A driver's effect on heterogeneity thus depends critically on when it is acquired. A strong driver acquired early expands the intermediate lineage and increases the supply of independent malignant founders, whereas the same driver acquired last strengthens the growth and age advantage of early-founded malignant clones. Under additive fitness effects, these opposing mechanisms can produce a non-monotone relationship between selective advantage and clonal concentration. Threshold-like non-additive fitness effects can generate highly concentrated malignant populations, while order-dependent terminal fitness causes the faster-growing pathway to dominate asymptotically. These results show how mutation order, mutational accessibility, selection, and epistasis jointly determine lineage-level intratumor heterogeneity.

q-bio.PE

Phase transitions in microbial lineage trees

Microbial populations exhibit high cell-to-cell variability, which fundamentally shapes population behavior. A striking consequence is the existence of phase transitions, where small genetic or environmental changes trigger abrupt shifts in population dynamics. While biological phase transitions have often been proposed, connecting observed behavior to the underlying physics has remained challenging. We combine population genetics with statistical physics to show how phase transitions arise naturally in microbial populations. We highlight the existence of a first-order transition in a model of bacterial plasmid engineering and find a strict lower bound on the number of plasmids that can be stably maintained in a population.

q-bio.PE

Phylogenetic Inference and the Stickiness of Fréchet Means, via Precise Asymptotics of an Embedded Random Walk

A well-known phenomenon in statistical analyses of populations of phylogenetic trees in the Billera-Holmes-Vogtmann space is that the topology of the Fréchet mean tree can contain multifurcations (i.e., internal nodes with more than two children), which raises the practical question of whether this reflects a population-level branching structure (hard polytomy) or merely sampling variability in the data (soft polytomy). This is an instance of the more general phenomenon of "stickiness" in non-Euclidean statistics, whereby the sample Fréchet mean in certain non-positively curved stratified spaces becomes permanently trapped in a lower-dimensional stratum. In this work, we identify a particular multidimensional random walk embedded within the Fréchet mean process, and we show that the time at which stickiness occurs is determined by the largest last-passage time above zero of the coordinates of this random walk. Using this representation, we develop a fully nonparametric procedure for estimating the probability that trifurcations in a sample Fréchet mean tree will bifurcate at some future time if more observations are collected. Lastly, we apply our methodology to a problem in phylogenetics where we consider whether an observed trifurcation in the species tree of primates, glires, and tree shrews is genuinely trifurcated at the population level.

q-bio.PE