Search arXiv⌕ Search

arXiv · 2512.08140

Non-parametric assessment of the calibration of individualized treatment effects

Abstract

An important aspect of the performance of algorithms that predict individualized treatment effects (ITE) is moderate calibration, i.e., the average treatment effect among individuals with predicted treatment effect of z being equal to z. The assessment of moderate calibration is challenging on two fronts: counterfactual responses are unobserved, and quantifying the conditional response function for models that generate continuous predicted values requires regularization. Perhaps because of these challenges, there is currently no inferential method for the null hypothesis that an ITE model is moderately calibrated in a population. In this work, we propose non-parametric methods for the assessment of moderate calibration of ITE models for binary outcomes using data from a randomized trial. These methods simultaneously resolve both challenges, resulting in novel graphical, numerical, and inferential methods for the assessment of moderate calibration. The key idea is to formulate a stochastic process for the cumulative prediction errors that obeys a functional central limit theorem, enabling the use of the properties of Brownian motion for asymptotic inference. We propose two approaches to construct this process from a sample: a conditional approach that relies on predicted risks (often an output of ITE models), and a marginal approach based on replacing the cumulative conditional moments with their marginal counterparts. Numerical simulations confirm the desirable properties of both approaches and their ability to detect miscalibration of different forms. We use a case study to provide suggestions on graphical presentation and the interpretation of results. Moderate calibration of predicted ITEs can be assessed without requiring regularization techniques or making assumptions about the functional form of treatment response. The accompanying cumulcalib R package implements this method.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mohsen Sadatsafavi, Jeroen Hoogland, Thomas P. A. Debray, John Petkau. 2026-08-18. Non-parametric assessment of the calibration of individualized treatment effects. https://doi.org/10.1002/sim.70724

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bayesian Neural-Net-Assisted Multi-Treatment Mixture Cure Survival Model with Application in Pediatric Oncology

Estimating covariate-conditional treatment effects in multi-arm oncology studies is complicated when treatment arms have common distributional features and a non-negligible fraction of patients achieve long-term remission. We propose a joint mixture cure model with covariate-dependent mixtures of log-normal kernels with treatment-specific inclusions. Both linear and neural-network-assisted non-linear covariate links are proposed. Specifically, the susceptible survival distributions use a common finite dictionary of log-normal components, and then a binary inclusion matrix determines which components are active in each treatment arm. All parameters, including the hidden bases of the neural network, are learned jointly, while the output coefficients remain treatment- or component-specific. Posterior inference is performed using gradient-based MCMC, and treatment effects are summarized by covariate-conditional differences in restricted mean survival time (RMST). Variable importance is assessed using thresholded marginal best linear projections with data partitioning. Across two simulation settings, the proposed method demonstrates good finite-sample performance, with lower RMST-based estimation error than flexsurvcure. Compared with pairwise grf fits, the proposed method yields lower RMST-contrast MSE in most comparisons while ensuring mutually coherent multi-treatment contrasts. Finally, the application to the AALL0434 trial reveals covariate-dependent patterns in RMST posterior across methotrexate-based regimens and provides new insights into how these differences vary with patient covariates, highlighting the method's practical utility for studying heterogeneous treatment effects in pediatric oncology trials.

stat.ME↗

A Three-Stage PCA Procedure for Sequentially Arriving High-Dimensional Data

We develop a three-stage adaptive procedure for principal component analysis (PCA) when high-dimensional observations are collected sequentially and additional sampling incurs a cost. The procedure balances PCA compression loss against sampling cost while selecting the retained dimension through a prescribed explained-variance criterion. Starting from a pilot sample, an intermediate stage updates the PCA quantities before determining the final sample size, thereby avoiding reliance on unknown population eigenvalues. Under suitable regularity conditions, we establish both first- and second-order efficiency relative to the population oracle. Comparison with the corresponding two-stage rule shows that the additional recalibration yields sharper second-order control and reduces the influence of the pilot stage on the final sampling decision. The theory allows the ambient dimension to exceed the sample size under appropriate covariance and spectral conditions. Simulation studies demonstrate the strong finite-sample performance of the procedure across increasing dimensions and several dense covariance structures. As a real-data application, we conduct a retrospective study of gene-expression data from 32 cancer-type cohorts in The Cancer Genome Atlas, illustrating both cost-effective early stopping and settings in which additional observations are recommended.

stat.ME↗

Batting Average as the Product of Two Rates: Skill, Luck, and the Disappearance of the .400 Hitter

Batting average factors exactly as BA = c times f, where c = (AB - SO)/AB is the rate of avoiding a strikeout and f = H/(AB - SO) is the rate at which non-strikeout at-bats become hits. Using the 256 major-league hitters with at least 300 at-bats in 2025, we show that the two factors behave very differently. Strikeout avoidance is highly repeatable, with a median year-to-year correlation of 0.86 over 19 consecutive-season pairs from 2004 to 2025. The finishing rate is not: its median correlation is 0.44, and batting average itself (0.44) is no more repeatable than its noisier factor. A bivariate logistic-normal random-effects model fit to the 2025 season estimates the correlation between the two talents at -0.68 (95\% profile interval -0.85 to -0.50), far stronger than the raw correlation of -0.44, and implies single-season reliabilities of 0.91 for c and 0.37 for f. The model also yields a closed-form bivariate shrinkage estimator in which a hitter's strikeout rate informs the estimate of his finishing rate. Applied to decades of American and National League data, the decomposition revisits Gould's explanation for the disappearance of the .400 hitter. Since the dead-ball era, the talent variance of strikeout avoidance has grown roughly 4.6-fold while that of finishing has halved, and the correlation between them has moved from near zero to about -0.53. That emerging trade-off, rather than a general narrowing of talent, accounts for the reduced spread of modern batting averages.

stat.ME↗