arXiv · 2609.23747
A Stochastic EM Algorithm with Sampling-Importance Resampling for Missing Data in Regression with Nonlinear Predictors
Abstract
Estimating regression models with nonlinear predictor transformations is challenging when data are missing, because nonlinearity typically renders the conditional distribution of the missing values intractable. Previous methods require specific nonlinear forms, such as polynomials or interactions, or rely on approximations that can induce bias. We propose a stochastic EM algorithm that uses sampling-importance resampling (SIR-StEM) to handle missing data under arbitrary nonlinear transformations of predictors, either substantively motivated or incidental, such as spline basis expansions. Unlike approaches that require linearity or closed-form conditionals, SIR-StEM only requires evaluating the complete-data likelihood up to proportionality, making it applicable across a broad class of nonlinear regression models. We construct an algorithm that makes use of missing data pattern information for computational efficiency, and establish asymptotic normality of the estimator. We demonstrate the method in two simulation studies, one with parametric nonlinear transformations and another with a spline basis expansion based on real behavioral measures. Results show that SIR-StEM yields low bias and near-nominal confidence interval coverage, outperforming other common approaches. We conclude with limitations and directions for future research.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dale S. Kim. 2026-09-20. A Stochastic EM Algorithm with Sampling-Importance Resampling for Missing Data in Regression with Nonlinear Predictors. https://arxiv.org/abs/2609.23747
Cite the original work for its findings. Save a collection to share your selection of sources.