Search arXiv⌕ Search

arXiv · 2609.37064

A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI

Abstract

AI models trained on the same data can disagree about patient risk, with disagreement potentially concentrated in clinically important subgroups. We propose V(S), a statistically grounded vulnerability index combining an observable lower-bound witness of model disagreement with clinical severity, and develop inference and multiplicity-adjustment procedures for auditing prespecified subgroups. We applied the framework to two large critical-care cohorts, MIMIC-IV (n = 65,078) for model development and eICU-CRD (n = 188,230 admissions, 208 hospitals) for external validation, comparing a random forest with logistic regression across 158 prespecified subgroups. The two primary models did not both satisfy the prespecified epsilon = 0.02 Rashomon-set tolerance: the logistic-regression AUC was 0.0488 below the best candidate-model AUC. The RF-LR discrimination gap is thus interpreted as disagreement between two specific models, not as a guaranteed lower bound on the full Rashomon set. Nine subgroups had discrimination gaps distinguishable from a prespecified clinical floor. The age >=80 and cardiac subgroup had the largest point estimate of V(S) (0.307), but was underpowered and did not meet the full high-priority decision rule. The univariate cardiac subgroup (V(S) = 0.193, 95% CI [0.163, 0.223]) was the only statistically distinguishable subgroup with adequate power. Post hoc analyses identified lactate as important for both models but did not establish a causal explanation for the disagreement. Four simulation studies quantified operating characteristics of the proposed procedures, including inflated small-sample detection rates and imperfect Wald-interval coverage. The framework offers a reproducible approach for ranking subgroup vulnerability to model disagreement while separating exploratory signals from adequately supported findings.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Enock Adu Bonsu. 2026-09-29. A statistical framework for identifying subgroup vulnerability to predictive multiplicity in clinical AI. https://arxiv.org/abs/2609.37064

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Delaunay Weighted Two-sample Test for High-dimensional Data by Incorporating Geometric Information

Two-sample hypothesis testing is a fundamental problem with various applications, which faces new challenges in the high-dimensional context. To mitigate the issue of the curse of dimensionality, high-dimensional data are typically assumed to lie on a low-dimensional manifold. To incorporate geometric information in the data, we propose to apply the Delaunay triangulation and develop the Delaunay weight to measure the geometric proximity among data points. In contrast to existing similarity measures that only utilize pairwise distances, the Delaunay weight can take both the distance and direction information into account. A detailed computation procedure is developed to learn the unknown manifold and approximate the Delaunay weight. We further propose a novel nonparametric test statistic using the Delaunay weight matrix. Asymptotic normality under the null and consistency under the alternative of the test statistic are developed. Applied to simulated data, the new test shows robustness to the learning of the unknown manifold and exhibits substantial power gain if the distributions differ in the principal directions of covariance matrices. The proposed test also detects significant differences on a real dataset of mice protein expression levels.

stat.ME↗

Randomization Tests in Switchback Experiments

Switchback experiments assign an experimental unit, such as a market or a platform, to treatment or control over successive blocks of time. Inference can be challenging because these experiments often contain only a small number of randomized blocks, while outcomes may exhibit serial dependence, seasonality, and treatment effects that persist across periods. We develop a conditional randomization test for the null of no total treatment effect that is finite-sample valid under standard assumptions on temporal interference without requiring a parametric model for outcomes. We also develop randomization tests for these assumptions: a carryover test that assesses whether past assignments continue to affect outcomes beyond a prespecified horizon and a non-anticipation test that assesses whether future assignments affect current outcomes. For hypotheses about average treatment effects, we establish asymptotic validity of studentized randomization tests under additional regularity conditions. Finally, we derive power approximations and characterize the tradeoffs between experimental design choices, the number of informative randomized comparisons, and statistical power. Numerical experiments with stylized potential outcomes and a dynamic rideshare model illustrate finite-sample performance and implications for experimental design.

stat.ME↗

Minimum Specification Perturbation: Robustness as Distance-to-Falsification in Causal Inference

Empirical causal claims depend on many analyst decisions, from selecting covariates to choosing estimators. Existing robustness tools summarize how results vary across these choices, but, to the best of our knowledge, do not answer: \textbf{How many analyst decisions must change to reach a specification, which is a set of choices, whose confidence interval (CI) contains zero?} We introduce \emph{Minimum Specification Perturbation (MSP)}, the smallest number of changes. MSP is small under the null, grows with effect strength and captures distance-to-falsification information that dispersion-based summaries cannot report; when making decisions under weak effects, an MSP-based rule yields lower false-positive rates than dispersion-based rules. We show that Fragility Index and MSP measure orthogonal vulnerabilities: fragility to influential observations need not imply fragility to specification choices. On the LaLonde benchmark, MSP = 1 implies that one decision change makes the CI contain zero. We further provide exact permutation calibration under randomization and characterize computation, showing tractable cases under additive structure and NP-hardness in general.

stat.ME↗