Search arXiv⌕ Search

arXiv · 0708.0169

Data-driven goodness-of-fit tests

Abstract

We propose and study a general method for construction of consistent statistical tests on the basis of possibly indirect, corrupted, or partially available observations. The class of tests devised in the paper contains Neyman's smooth tests, data-driven score tests, and some types of multi-sample tests as basic examples. Our tests are data-driven and are additionally incorporated with model selection rules. The method allows to use a wide class of model selection rules that are based on the penalization idea. In particular, many of the optimal penalties, derived in statistical literature, can be used in our tests. We establish the behavior of model selection rules and data-driven tests under both the null hypothesis and the alternative hypothesis, derive an explicit detectability rule for alternative hypotheses, and prove a master consistency theorem for the tests from the class. The paper shows that the tests are applicable to a wide range of problems, including hypothesis testing in statistical inverse problems, multi-sample problems, and nonparametric hypothesis testing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mikhail Langovoy. 2017-09-21. Data-driven goodness-of-fit tests. https://arxiv.org/abs/0708.0169

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Optimal Transport Based Testing in Factorial Designs

We introduce a general framework for testing statistical hypotheses in factorial designs for probability measures supported on discrete spaces. The suggested methodology is based on the pairwise comparison of measures using optimal transport (OT). The formulation of hypotheses is intuitive: It is a direct extension of those underlying the analysis of variance (ANOVA) and its nonparametric counterparts to test for linear relationships between (discrete) probability measures in factorial designs. To this end, means or cumulative distribution functions simply will be replaced by measures. We derive under the null hypotheses and under (local) alternatives the asymptotic distribution of the corresponding empirical OT test statistic, which is the optimal value of a linear program with random objective function. It turns out that this requires to extend existing techniques from probability measures to signed measures, and we show directional Hadamard differentiability and the validity of the functional delta method. We discuss computational issues, permutation and bootstrap tests, and back up our findings with simulations. We illustrate our methodology on datasets from cellular biophysics and from biometric fingerprint identification.

math.ST↗

Stable Central Limit Theorems for Discrete-Time Lag Martingale Difference Arrays: Applications to Dynamic Causal Inference

Recent work in dynamic causal inference introduced a class of discrete-time stochastic processes that generalize martingale difference sequences and arrays as follows: the random variates in each sequence have expectation zero given certain lagged filtrations but not given the natural filtration. We formalize this class of stochastic processes and prove stable central limit theorems (CLTs) via martingale-coboundary decomposition, leveraging the classical martingale CLT. We develop a variety of sufficient conditions, including conditions under which the limiting variance has a simple form that depends on variances and covariances of neighboring variates. We demonstrate the application of these results to inference for time-averaged treatment effects in switchback designs and present a simulation study supporting their validity. The CLTs enable various extensions to existing methodology for design-based approaches to dynamic causal inference, including time-lagged effects, random limiting variances, cross-unit dependence, and vector-valued estimands.

math.ST↗

Exact Likelihood-Coin Poisson Sampling for Bayesian Inverse Problems with Sharp Complexity Bounds

We develop an exact posterior-sampling framework for Bayesian inverse problems when selected bounded forward observables can be accessed through Bernoulli events. A Bernstein--Poisson construction converts these forward coins into scaled Gaussian likelihood coins, and thinning an inflated prior Poisson point process yields posterior atoms that are iid conditional on their number. For independent Gaussian observations, we derive an exact mean-work identity for the implemented early-stopped factory and sharp small-noise complexity laws governed by local prior-predictive mass near the exact-fit set; a factorial-moment construction extends the likelihood factory to correlated Gaussian errors. The posterior algorithm is model-agnostic once Bernoulli access is available. As one continuum realization, we use Feynman--Kac sampling, Poisson killing, and lazy random-series evaluation for a bounded elliptic resolvent problem with a function-valued coefficient. Numerical experiments validate the forward and likelihood coins, support the predicted work regimes, and demonstrate posterior sampling without deterministic spatial discretization or fixed parameter truncation in the target. A matched-accuracy benchmark against finite-difference prior rejection illustrates how deterministic discretization bias changes the posterior accuracy--cost balance.

math.ST↗