Search arXiv⌕ Search

arXiv · 2508.19070

Replicability: Terminology, Measuring Success, and Strategy

Abstract

Empirical science needs to be based on facts and claims that can be reproduced. This calls for replicating the studies that proclaim the claims, but practice in most fields still fails to implement this idea. When such studies emerged in the past decade, the results were generally disappointing. There have been an overwhelming number of papers addressing the ``reproducibility crisis'' in the last 20 years. Nevertheless, terminology is not yet settled, and there is no consensus about when a replication should be called successful. This paper intends to clarify such issues. A fundamental problem in empirical science is that usual claims only state that effects are non-zero, and such statements are scientifically void. An effect must have a \emph{relevant} size to become a reasonable item of knowledge. Therefore, estimation of an effect, with an indication of precision, forms a substantial scientific task, whereas testing it against zero does not. A relevant effect is one that is shown to exceed a relevance threshold. This paradigm has implications for the judgement on replication success. A further issue is the unavoidable variability between studies, called heterogeneity in meta-analysis. Therefore, it is of little value, again, to test for zero difference between an original effect and its replication, but exceedance of a corresponding relevance threshold should be tested. In order to estimate the degree of heterogeneity, more than one replication is needed, and an appropriate indication of the precision of an estimated effect requires such an estimate. These insights, which are discussed in the paper, show the complexity of obtaining solid scientific results, implying the need for a strategy to make replication happen.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Werner A. Stahel. 2025-08-26. Replicability: Terminology, Measuring Success, and Strategy. https://arxiv.org/abs/2508.19070

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Asymptotic confidence intervals for the difference and the ratio of the weighted kappa coefficients of two diagnostic tests subject to a paired design

The weighted kappa coefficient of a binary diagnostic test is a measure of the beyond-chance agreement between the diagnostic test and the gold standard, and depends on the sensitivity and specificity of the diagnostic test, on the disease prevalence and on the relative importance between the false positives and the false negatives. This article studies the comparison of the weighted kappa coefficients of two binary diagnostic tests subject to a paired design through confidence intervals. Three asymptotic confidence intervals are studied for the difference between the parameters and five other intervals for the ratio. Simulation experiments were carried out to study the coverage probabilities and the average lengths of the intervals, giving some general rules for application. A method is also proposed to calculate the sample size necessary to compare the two weighted kappa coefficients through a confidence interval. A program in R has been written to solve the problem studied and it is available as supplementary material. The results were applied to a real example of the diagnosis of malaria.

stat.OT↗

Engaging students with statistics through choice of real data context on homework

Statistics educators recommend teaching with real data with relevant contexts, but defining relevancy is challenging and varies by student. We investigated whether providing student choice of data context increases engagement through a quasi-experiment in two sections of an introductory probability and statistics course at a large public university (n=65 consenting students). Sections alternated as treatment and control: during their treatment, students chose weekly homework from three similar instructor-provided options varying by data context; during control weeks, they received randomly assigned contexts. We found no significant difference in homework grades between treatment and control conditions. However, thematic analysis revealed students with choice reported enhanced engagement and motivation, greater appreciation for statistics' real-world value, and increased autonomy. Students overwhelmingly preferred contexts relevant to their interests, experiences, daily lives, and career paths-though preferences varied considerably across individuals. Based on these findings, we provide four recommendations for statistics educators: (1) use real data with authentic contexts, (2) select contexts students care about, (3) incorporate variety across data contexts, and (4) consider choice as a pedagogical tool.

stat.OT↗

See You at the Posterior Line: Learning Bayesian Modeling Through a Car Racing Game

We present an interactive classroom activity designed to address a central challenge in teaching introductory Bayesian statistics: how to formalize subjective knowledge and available information into prior distributions and then update them with empirical data. Role-playing as data analysts for a racing team, students evaluate candidate tires by converting qualitative engineering reports into prior distributions, collecting primary data via a virtual racing game, and using a Beta-Binomial model to inform team strategy. This discovery-based exercise allows small groups to observe directly how different prior choices and sample data jointly shape posterior inference. Student feedback ($n=32$) highlights high enjoyment, engagement and improved conceptual clarity. Open-access materials to implement the activity are provided, alongside recommendations for adapting it to other teaching contexts.

stat.OT↗