Search arXivSearch

arXiv · 2603.03591

Mutation Rate Variation Across Genomic Regions in \textit{Arabidopsis thaliana}

Abstract

In population genetics, mutation rate is often treated as a homogeneous parameter across the genome. Empirical evidence, however, shows systematic variation across genomic contexts associated with chromatin organization and epigenomic features. Using gene-level de novo mutation data from Arabidopsis thaliana, we test whether chromatin features predict not only the mean per-base mutation rate but also its variability across genes. To reduce heterogeneity in selective regime, we restrict analysis to essential and lethal loci subject to strong purifying selection. Across complementary multivariable models including heteroskedasticity-robust linear regression, length-weighted regression, and Poisson generalized linear models with exposure offsets, histone marks associated with active transcription (H3K4me1, H3K4me3, H3K36ac) are consistently associated with lower mean mutation rates and substantially reduced between-gene variance. GC content shows little association with the mean once chromatin predictors are controlled but is positively associated with mutation-rate variability. Estimates of skewness and kurtosis reveal no significant higher-order structure attributable to epigenomic predictors. A standardized Tajima's $D$ statistic yields directionally consistent but statistically underpowered associations with both the mean and variance of gene-level mutation rates. These results indicate that mutation rate is systematically structured by chromatin state within functionally constrained genes and suggest that evolutionary processes may act not only on expected mutation rate but also on its variability across loci.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Elisa Heinrich-Mora, Marcus W. Feldman. 2026-03-03. Mutation Rate Variation Across Genomic Regions in \textit{Arabidopsis thaliana}. https://arxiv.org/abs/2603.03591

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Beta-coalescents when sample size is large

Sweepstakes reproduction refers to a highly skewed individual recruitment success without involving natural selection and may apply to individuals in broadcast spawning populations characterised by Type III survivorship. We consider an extension of the model of sweepstakes reproduction for a haploid panmictic population of constant size $N$; the extension also works as an alternative to the Wright-Fisher model. Our model incorporates an upper bound on the random number of potential offspring (juveniles) produced by a given individual. Depending on how the bound behaves relative to the total population size, we obtain the Kingman coalescent, an incomplete Beta-coalescent, or the (complete) Beta-coalescent. We argue that applying such an upper bound is biologically reasonable. Moreover, we estimate the error of the coalescent approximation. The error estimates reveal that convergence can be slow, and small sample size can be sufficient to invalidate convergence, for example if the stated bound is of the form $N/\log N$. We use simulations to investigate the effect of increasing sample size on the site-frequency spectrum. When the limit is a Beta-coalescent, the site frequency spectrum will be as predicted by the limiting tree even though the full coalescent tree may deviate from the limiting one. When in the domain of attraction of the Kingman coalescent the effect of increasing sample size depends on the effective population size as has been noted in the case of the Wright-Fisher model. Conditioning on the population ancestry (the random ancestral relations of the entire population at all times) may have little effect on the site-frequency spectrum for the models considered here (as evidenced by simulation results).

q-bio.PE

The role of nestedness and saturating feedback in bipartite ecological systems

Large ecosystems balance competition and cooperation, yet standard generalized Lotka--Volterra models make mutualism destabilizing by amplifying disorder and driving unbounded growth. We show that Monod-like saturation resolves this paradox: dynamical mean-field theory and random-matrix analysis reveal a broader stable phase and enhanced survival. Network architecture provides a second control mechanism, but nestedness offers no intrinsic stability advantage. Instead, it is a byproduct of degree distributions with high connectivity necessary for stability.

q-bio.PE

TreeFlow: probabilistic modelling and automatic differentiation for phylogenetics

Probabilistic modelling frameworks are powerful tools for statistical modelling and inference. They are not immediately generalizable to phylogenetic problems due to the particular computational properties of the phylogenetic tree object. TreeFlow is a software library for probabilistic modelling and automatic differentiation with phylogenetic trees. It embeds phylogenetic trees in the TensorFlow Probability framework, and implements inference algorithms for phylogenetic models given a fixed tree topology. We demonstrate how TreeFlow can be used to quickly implement and assess new models. We also show that it provides reasonable performance for gradient-based inference algorithms compared to specialized computational libraries for phylogenetics.

q-bio.PE