Search arXiv⌕ Search

arXiv · 1807.05996

Lectures on Statistics in Theory: Prelude to Statistics in Practice

Abstract

This is a writeup of lectures on "statistics" that have evolved from the initial version for the 2009 Hadron Collider Physics Summer School at CERN to versions for other venues and, most recently, for the African School of Fundamental Physics and Applications in 2024. The emphasis is on foundations, using simple examples to illustrate the points that are still debated in the professional statistics literature. The three main approaches to interval estimation (Neyman confidence, Bayesian, likelihood ratio) are discussed and compared in detail, with and without nuisance parameters. Hypothesis testing is discussed mainly from the frequentist point of view, with pointers to the Bayesian literature. Various foundational issues are emphasized, including the conditionality principle and the likelihood principle.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert D. Cousins. 2024-06-26. Lectures on Statistics in Theory: Prelude to Statistics in Practice. https://arxiv.org/abs/1807.05996

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Statistical validation of calorimeter inpainting with generative diffusion priors

Localized detector inefficiencies produce incomplete calorimeter data that limit the ability to perform precision measurements. We address this problem in relativistic heavy-ion collisions from a Bayesian perspective using pretrained calorimeter diffusion models as priors to reconstruct the missing signal conditioned on surrounding measurements. In this work, we conduct a systematic comparison of several diffusion-based inpainting algorithms, whose performance is evaluated using Bayesian posterior diagnostics of energy response, spatial bias, and uncertainty calibration. The reconstruction fidelity is also analyzed across collision centralities and masked region sizes. This study establishes a general validation strategy for probabilistic reconstruction of missing detector information.

physics.data-an↗

Attributing extreme-event probability to a source variable

Information-flow theory quantifies directional coupling through the rate of change of a target's Shannon entropy, a bulk functional insensitive to the tail of the distribution. We ask instead how a source variable contributes to the probability that the target exceeds a threshold. For source-additive drift the source's share of the marginal probability current is exact, and it separates into a mean-forcing part and a conditional-excess part that vanishes under independence. An identity links the two descriptions: the Liang information flow is the density-weighted mean of the derivative of the specific source current, whereas the exceedance current is its level at the threshold. This explains why entropy-based coupling collapses in saturated regimes where the source contributes most to the extreme; the Rényi information flow, a higher-order expansion and a Fisher-normalised response fail for the same reason. The flux decomposition, although exact, does not attribute: its terms are gross transports that nearly cancel. The quantity that does attribute is an adjoint response built from the backward generator which, read as a relative change, is uniformly accurate across two decades of event probability. We map the operating envelope of the estimator under omitted drivers, hidden slow memory, state-dependent coupling and multiplicative noise. Under correct specification the attribution carries a reproducible shortfall of ten to twenty-five per cent; the misspecifications we test bias it upward by up to ninety per cent. Applied to the 2003 European and 2010 Russian heatwaves in reanalysis, the attributed contribution of soil moisture is strongly threshold dependent, rising from the per cent level at moderate thresholds to a factor of two at the rarest, so a contribution quoted without its threshold is underspecified.

physics.data-an↗

Simulation-Based Inference and Unbinned Asimov Construction with Hybrid Neural Density Estimation

High-dimensional, unbinned neural simulation-based inference often relies on neural ratio estimation, which uses expressive supervised models to estimate density ratios, but a learned ratio by itself provides neither an explicit normalized density nor a generative model. Flow-based surrogate models instead enable tractable density evaluation and efficient sampling, but residual density-estimation errors can limit precision for complex implicit distributions. We propose \textit{hybrid neural density estimation}, which uses a flow to define a parameter-independent reference distribution, and classifiers to estimate target-to-reference density ratios. Multiplying a learned ratio by the reference density gives an evaluable target density surrogate. The ratio also provides importance weights for integration and resampling. We show how this representation defines an exact Asimov dataset, whose maximum likelihood fit returns the generating parameters. We also show how the tractable reference supplies renewable samples for pseudo-experiments and how it enables methods to reduce the Monte Carlo variance of expected test statistic calculations. We demonstrate the construction in a toy statistical model motivated by high-energy physics measurements, but for which the exact densities are known analytically.

physics.data-an↗