A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI
AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference and multiplicity-adjustment procedures for auditing prespecified subgroups. We applied the framework to two large critical-care cohorts, MIMIC-IV (n = 65,078) for model development and eICU-CRD (n = 188,230 admissions, 208 hospitals) for external validation, comparing a random forest with logistic regression across 158 prespecified subgroups. The two primary models did not both satisfy the prespecified epsilon = 0.02 Rashomon-set tolerance: the logistic-regression AUC was 0.0488 below the best candidate-model AUC. The RF-LR discrimination gap is thus interpreted as disagreement between two specific models, not as a guaranteed lower bound on the full Rashomon set. Nine subgroups had discrimination gaps distinguishable from a prespecified clinical floor. The age >=80 and cardiac subgroup had the largest point estimate of V(S) (0.307), but was underpowered and did not meet the full high-priority decision rule. The univariate cardiac subgroup (V(S) = 0.193, 95% CI [0.163, 0.223]) was the only statistically distinguishable subgroup with adequate power. Post hoc analyses identified lactate as important for both models but did not establish a causal explanation for the disagreement. Four simulation studies quantified operating characteristics of the proposed procedures, including inflated small-sample detection rates and imperfect Wald-interval coverage. The framework offers a reproducible approach for ranking subgroup vulnerability to model disagreement while separating exploratory signals from adequately supported findings.