Finite-Sample FDR Control for Greedy Aggregation over Networks
Distributed multiple testing asks $N$ sites, each holding p-values for its own hypotheses, to control the false discovery rate (FDR) of the discoveries made across the whole network while communicating only a small number of bits. Greedy interval aggregation (Pournaderi and Xiang, IEEE TSIPN, 2023) meets the communication budget but controls FDR only asymptotically, and its FDR can exceed the target at finite sample sizes; the same interval counts both select the rejection regions and calibrate the stopping rule, which biases the selected interval densities upward. We propose \emph{budgeted BONuS-GA}: each site mixes synthetic uniform p-values into its data, the center ranks candidate p-value intervals across sites using the pooled counts, and each site's false discoveries are estimated from its own synthetic counts under a per-site share of the level. The procedure needs no knowledge of the null proportions, uses every p-value for both selection and inference, and controls $FDR\leα$ at every sample size for any fixed assignment of hypotheses to sites, assuming only that the null p-values are independent uniforms, independent of the non-null p-values. It keeps the original $O(\sqrt m\log m)$ communication, $m$ being the total number of p-values, when $N=O(\sqrt m)$. We also give a sample-splitting alternative, cross-fit greedy aggregation, with finite-sample control under a random-site model, and we quantify the selection bias behind the original procedure's failure. In a monitoring-network simulation the budgeted procedure retains most of the power of its heuristic counterpart at moderate-to-large site sizes.