Search arXivSearch

arXiv · 2606.03384

Evolution as a Process of Causal Inference

Abstract

Recently, the mapping of the replicator equation onto Bayes' theorem has been recognised, leading to an analogy between evolutionary dynamics and Bayesian learning. However, this analogy holds only for pure selection in infinite populations and breaks down when mutations -- a central mechanism of evolution -- are introduced. Here I propose that evolution by natural selection, at least for populations of haploid replicators in static environments, is best understood not as a learning process but as a process of causal inference. Each mutation event constitutes a natural experiment in which the parent serves as the control and the mutant offspring as the treated unit. Natural selection screens the causal effect of the mutation on fitness, retaining mutations with non-negative effects. I formalise this view within the Neyman-Rubin potential-outcomes framework. I first develop the general theory using a generic fitness outcome and show how the core identification assumptions in causal inference (Stable Unit Treatment Value Assumption, Consistency, Unconfoundedness, Positivity) map onto evolutionary biology. Using the unnormalised quasispecies equation, I prove that the intergenerational change in mean fitness decomposes exactly into a selection term -- recovering Fisher's Fundamental Theorem -- plus a mutation term that corresponds to a fitness-weighted average of the cumulated effect of all mutations over all parental genotypes. I show that this decomposition extends, under suitable assumptions, to the generalised replicator-mutator equation and that the frequencies of populations of matched parents-offspring update in proportion to the average causal effect of mutations on fitness.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jacopo Iacovacci. 2026-06-02. Evolution as a Process of Causal Inference. https://arxiv.org/abs/2606.03384

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Anomalous First Passage in Evolution: Edge-KPZ Theory

The pace of evolution depends on how rapidly new phenotypes arise. We show that neutral Wright-Fisher evolution exhibits anomalous first passage despite diffusive mutations. The mean time for the first individual to reach a prescribed phenotypic distance scales approximately as $(σ^2)^{-3/2}$ with mutation variance $σ^2$. Two crossovers bound this regime, with inverse-variance scaling on either side. Combining coalescent theory with Kardar-Parisi-Zhang (KPZ) fluctuations at the dilute population edge, we develop an edge-KPZ theory of all three regimes. The anomaly persists under weak selection.

q-bio.PE

Graph construction in QUBO-based recursive phylogenetic tree reconstruction

Molecular sequence data are used to reconstruct evolutionary relationships among taxa, but reconstruction accuracy depends not only on the tree-building method but also on how pairwise sequence relationships are represented. We evaluated sequence-to-affinity representations in a recursive normalized-cut (Ncut) framework whose graph-partitioning subproblems were formulated as quadratic unconstrained binary optimization (QUBO) models and solved using Simulated Bifurcation. Using simulated amino-acid and nucleotide datasets spanning multiple tree-generation settings and evolutionary divergence, we compared normalized bit-score affinities with representations derived from transformed sequence similarities and evolutionary distances, examined post-swap refinement, and used neighbor joining (NJ) as a distance-based comparator. Affinity representation substantially affected internal split recovery, particularly for nucleotide data. JC69-based local affinities maintained comparatively high accuracy as divergence increased, whereas normalized bit-score and BLAST-derived kernel representations declined more markedly. Post-swap refinement generally improved recovery, but not consistently across individual reconstructions. NJ achieved higher mean split recovery than corresponding recursive Ncut reconstructions for WAG and JC69 distances across all evaluated conditions, whereas recursive Ncut outperformed NJ for BLAST-derived logarithmic distances under some conditions. These results show that graph construction is an important determinant of recursive Ncut-based phylogenetic reconstruction. A representation that performs well within Ncut does not necessarily provide the most accurate use of the underlying pairwise distances. Pairwise representation, affinity transformation, optimization, and recursive tree construction should therefore be evaluated jointly.

q-bio.PE

Beta-coalescents when sample size is large

Sweepstakes reproduction refers to a highly skewed individual recruitment success without involving natural selection and may apply to individuals in broadcast spawning populations characterised by Type III survivorship. We consider an extension of the model of sweepstakes reproduction for a haploid panmictic population of constant size $N$; the extension also works as an alternative to the Wright-Fisher model. Our model incorporates an upper bound on the random number of potential offspring (juveniles) produced by a given individual. Depending on how the bound behaves relative to the total population size, we obtain the Kingman coalescent, an incomplete Beta-coalescent, or the (complete) Beta-coalescent. We argue that applying such an upper bound is biologically reasonable. Moreover, we estimate the error of the coalescent approximation. The error estimates reveal that convergence can be slow, and small sample size can be sufficient to invalidate convergence, for example if the stated bound is of the form $N/\log N$. We use simulations to investigate the effect of increasing sample size on the site-frequency spectrum. When the limit is a Beta-coalescent, the site frequency spectrum will be as predicted by the limiting tree even though the full coalescent tree may deviate from the limiting one. When in the domain of attraction of the Kingman coalescent the effect of increasing sample size depends on the effective population size as has been noted in the case of the Wright-Fisher model. Conditioning on the population ancestry (the random ancestral relations of the entire population at all times) may have little effect on the site-frequency spectrum for the models considered here (as evidenced by simulation results).

q-bio.PE