Search arXiv⌕ Search

arXiv · 2610.07831

Design-Driven Inference for Online Experiments: Jointly Optimal Evidence Collection, Testing, and Estimation

Abstract

Online multifactorial experiments requires early stopping for practically negligible effects while controlling false stopping under continuous monitoring and retaining precise treatment-effect estimation. We develop a design-driven framework that chooses allocation and the evidence rule jointly, treating allocation as evidence collection and testing and estimation as complementary uses of the same evidence. For multifactor experiments with nuisance block effects and treatment-by-block interactions, we show that block-orthogonal designs remove nuisance contamination from the treatment score and form a complete class for worst-case evidence growth. Under a fixed information budget, an isotropic allocation and an explicit radial e-value jointly attain the minimax rate optimal for directionally unknown alternatives. The same allocation is universally optimal for maximum likelihood estimation of treatment effects, simultaneously achieving A-, D-, and E-optimality and establishing double optimality for testing and estimation. Batchwise replication yields an anytime-valid e-process, retains minimax rate optimality for cumulative conditional worst-case evidence growth, and preserves universal estimation optimality for the prespecified complete experiment. Simulations and a large language model prompt experiment illustrate gains in stopping efficiency and estimation precision.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jun Yu, Wenbiao, Zhao, Lixing, Zhu. 2026-10-06. Design-Driven Inference for Online Experiments: Jointly Optimal Evidence Collection, Testing, and Estimation. https://arxiv.org/abs/2610.07831

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Robust Bayesian methods using amortized simulation-based inference

Bayesian simulation-based inference (SBI) methods are used in statistical models where simulation is feasible but the likelihood is intractable. Standard SBI methods can perform poorly in cases of model misspecification, and there has been much recent work on modified SBI approaches which are robust to misspecified likelihoods. However, less attention has been given to the issue of inappropriate prior specification, which is the focus of this work. In conventional Bayesian modelling, there will often be a wide range of prior distributions consistent with limited prior knowledge expressed by an expert. Choosing a single prior can lead to an inappropriate choice, possibly conflicting with the likelihood information. Robust Bayesian methods, where a class of priors is considered instead of a single prior, can address this issue. For each density in the prior class, a posterior can be computed, and the range of the resulting inferences is informative about posterior sensitivity to the prior imprecision. We consider density ratio classes for the prior and implement robust Bayesian SBI using amortized neural methods developed recently in the literature. We also discuss methods for checking for conflict between a density ratio class of priors and the likelihood, and sequential updating methods for examining conflict between different groups of summary statistics. The methods are illustrated for several simulated and real examples.

stat.ME↗

On the permutation equivariance principle for causal estimands

In many causal inference problems, multiple action variables share a common causal role yet lack a natural ordering. \revblue{We consider $K\geq2$ action variables, each evaluated under treatment or control conditions,} and formalize permutation equivariance, the principle that permuting the variables permutes the corresponding estimands in a trackable manner, hence preserving their scientific meaning. We characterize this principle algebraically and present a complete class of weighted permutation equivariant estimands capturing main effects and interactions of all orders. We discuss the interpretation and choice of weights and characterize residual-free estimands, whose inclusion--exclusion sum recovers the endpoint contrast between the all-treated and all-control configurations. \revblue{Applying our general framework to network interference yields a new hierarchy of direct, indirect, and overall effects, whose first-order aggregates recover the average effects of \citet{hu2022average}. We also identify the overall effects as mixed derivatives of expected welfare under independent Bernoulli assignment.} We illustrate the framework through the contexts of factorial studies, causal mediation, and network interference.

stat.ME↗

Testing the Deviation of a High-dimensional Mean

This paper investigates the high-dimensional one-sample equivalence testing problem for a reference mean vector $\bbmu_0$: $H_0:\|\bbmu-\bbmu_0\|>d_0$ versus $H_1: \|\bbmu-\bbmu_0\|\leq d_0$ for a pre-specified length threshold $d_0 > 0$. In high-dimensional regimes, classical equivalence tests suffer from severe asymptotic breakdowns. To resolve this difficulty, we propose a novel test statistic based on a two-armed bandit (TAB) U-process, which leverages positive/negative feedback mechanisms from control theory. We rigorously establish the weak convergence of TAB U-process to a novel two-dimensional stochastic differential equation (SDE), termed the U-bandit SDE. By employing in-depth analysis of this SDE, we characterize the asymptotic size and power of the resulting test statistic. The theoretical framework is extended to the two-sample setting and validated through finite-sample simulations and an empirical analysis of microbiome compositions.

stat.ME↗