When is statistical evidence strong enough? Using hypothesis tests to value data collection
Policy decisions often hinge on conventional p-value thresholds, which ignore economic costs and benefits of further data collection. This paper recasts statistical significance as a choice between making an immediate policy recommendation and deferring it until further evidence is collected. The welfare-optimal decision corresponds, under minimax regret, to a statistical test whose level depends on the cost and precision of additional evidence. Inverting this rule, we introduce and recommend reporting the abstention-value (A-value) to determine where additional data collection is most needed. The A-value defines the break-even welfare cost of abstaining and recommending further experimentation given the initial evidence. Computing the A-value only requires a point estimate and its standard error, imposes no prior assumptions on policy effects, and can be converted to a monetary research budget using inputs already frequently used by regulators and funding agencies. When experimentation capacity is limited, we show that prioritizing additional data collection where A-values are the largest yields strong finite-sample guarantees. Applications to anti-poverty programs and to a medical study illustrate the practical benefits of using A-values.