Search arXivSearch

arXiv · 2602.05938

DiPPER: A Bayesian approach to differential prevalence analysis with applications in microbiome studies

Abstract

Recent evidence suggests that analyzing the presence/absence of taxonomic features can offer a compelling alternative to differential abundance analysis in microbiome studies. However, standard approaches to differential prevalence analysis face challenges with boundary cases and multiple testing. To address these limitations, we developed DiPPER (Differential Prevalence via Probabilistic Estimation in R), a method based on Bayesian hierarchical modeling. We benchmarked our method against existing differential prevalence methods, along with two differential abundance tools, using publicly available data from 57 human gut microbiome studies. We observed considerable variation in performance across the evaluated methods. Importantly, DiPPER demonstrated high sensitivity to detect potentially differentially prevalent features while maintaining a well-calibrated family-wise error rate under the global null hypothesis. Most notably, it outperformed the alternatives in the replication of findings across independent studies. Furthermore, DiPPER provides differential prevalence estimates and uncertainty intervals that are inherently adjusted for multiple testing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Juho Pelto, Kari Auranen, Janne V. Kujala, Leo Lahti. 2026-05-25. DiPPER: A Bayesian approach to differential prevalence analysis with applications in microbiome studies. https://arxiv.org/abs/2602.05938

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Switchback Experiments under Geometric Mixing

The switchback is an experimental design that measures treatment effects by repeatedly turning an intervention on and off for a whole system. Switchback experiments are a robust way to overcome cross-unit spillover effects; however, they are vulnerable to bias from temporal carryovers. In this paper, we consider properties of switchback experiments in Markovian systems that mix at a geometric rate. We find that, in this setting, standard switchback designs suffer considerably from carryover bias: Their estimation error decays as $T^{-1/3}$ in terms of the experiment horizon $T$, whereas in the absence of carryovers a faster rate of $T^{-1/2}$ would have been possible. We also show, however, that judicious use of burn-in periods can considerably improve the situation, and enables errors that decay as $\log(T)^{1/2}T^{-1/2}$. Our formal results are mirrored in an empirical evaluation.

stat.ME

Spectral divide-and-conquer MCMC for long stationary time series

The temporal dependence inherent in time series models poses a fundamental challenge for distributed Bayesian inference, as it precludes embarrassingly parallel algorithms based on naive independence assumptions in the time domain. We propose a frequency-domain framework for scalable Bayesian inference in stationary time series that exploits the asymptotic independence underlying the Whittle likelihood. To exploit parallel computing resources, we develop a distributed fast Fourier transform and integrate it with embarrassingly parallel Markov chain Monte Carlo (MCMC) algorithms within a modern cluster-computing framework. This enables the analysis of time series that exceed the memory capacity of a single computational node or for which computation time is a bottleneck. The proposed methodology is compatible with a broad class of existing divide-and-conquer algorithms for independent data by applying them to frequency-domain rather than time-domain partitions. We establish that the error of the spectral divide-and-conquer MCMC posterior approximation relative to the exact time-domain posterior converges to zero in probability in a shrinking neighbourhood of the full-data Whittle posterior mode. The corresponding convergence rate is also derived. Across several experiments, we demonstrate that our approach provides accurate approximations to the full-data Whittle posterior. The proposed method is shown to outperform the current state-of-the-art time-domain divide-and-conquer methodology, particularly for highly persistent processes. The methodology is further illustrated by fitting a semi-long range model to a long meteorological time series.

stat.ME

Robust filtering and smoothing via perturbation methods

Using a perturbation technique, we derive a new approximate filtering and smoothing methodology generalizing along different directions several existing approaches to robust filtering based on the score and the Hessian matrix of the observation density. The main advantages of the methodology can be summarized as follows: (i) it relaxes the critical assumption of a Gaussian conditional distribution for the latent states underlying such approaches; (ii) can be applied to a general class of state-space models including location, scale and count data models; (iii) rationalizes the approximation to the likelihood function within the same perturbation approach, thus allowing for straightforward inference of the model parameters; (iv) enables the computation of confidence bands around the state estimates reflecting the combination of parameter and filtering uncertainty. We show through an extensive Monte Carlo study that the mean square loss with respect to exact simulation-based methods is small in a wide range of scenarios. We finally illustrate empirically the application of the methodology to the estimation of stochastic volatility and correlations in financial time-series.

stat.ME