Search arXivSearch

arXiv · 2512.03235

Estimation of Semiparametric Factor Models with Missing Data

Abstract

We study semiparametric factor models in high-dimensional panels where the factor loadings consist of a nonparametric component explained by observed covariates and an idiosyncratic component capturing unobserved heterogeneity. A key challenge in empirical applications is the presence of missing observations, which can distort both factor recovery and loading estimation. To address this issue, we develop a projected principal component analysis (PPCA) procedure that accommodates general missing-at-random mechanisms through inverse-probability weighting. We establish consistency and derive the asymptotic distributions of the estimated factors and loading functions, allowing the sieve dimension to diverge and permitting the time dimension to be either fixed or growing. Unlike classical PCA, PPCA achieves consistent factor estimation even when T is fixed, and the limiting distributions under missing data exhibit mixture normality with enlarged asymptotic variances. Theoretical results are supported by simulations and an empirical application. Our findings demonstrate that PPCA provides an effective and robust framework for estimating semiparametric factor models in the presence of missing data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sijie Zheng. 2025-12-05. Estimation of Semiparametric Factor Models with Missing Data. https://arxiv.org/abs/2512.03235

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Unbiased Treatment Effect Estimation under Network Interference via Neighborhood-Excluded Cross-Fitting

Without interference, cross-fitting enables flexible covariate adjustment while preserving finite-sample unbiasedness under independent unit-level randomization. Under network interference, out-of-sample prediction alone no longer guarantees unbiasedness: assignments entering evaluation-fold Horvitz--Thompson weights may also affect outcomes in the training sample, inducing dependence between fitted predictions and those weights. We develop neighborhood-excluded cross-fitting, which constructs estimand- and design-specific training samples to restore the conditional independence needed for finite-sample unbiasedness without a correctly specified outcome model. We establish asymptotically valid design-based Wald inference for direct and indirect effects under Bernoulli randomization and for the global average treatment effect under Bernoulli cluster randomization. Neighborhood exclusion creates a trade-off in choosing the number of folds: unit-level splitting may require the number of folds to grow with average exclusion-neighborhood size, while cluster-level splitting can substantially relaxes this requirement, permitting a fixed number of folds under partial interference. For linear adjustment under Bernoulli randomization, we derive variance-optimal and confidence-interval-length-optimal procedures, establish explicit rate conditions allowing the covariate dimension to diverge, and show that the variance-optimal procedure is asymptotically no-harm. Simulations illustrate the bias from omitting neighborhood exclusion and the precision gains from adjustment. An application to a social network experiment yields confidence intervals substantially shorter than those from the unadjusted estimator.

stat.ME

Efficient estimation of weighted treatment effects under two-phase sampling

Two-phase sampling offers a practical way to collect costly confounders only in a subsample while retaining inexpensive information for a larger cohort. In observational causal studies, however, phase-2 selection can distort estimation of population causal effects if the sampling mechanism is ignored, and available phase-1 information may also be exploited to improve efficiency. Yet efficiency theory for causal estimands under such designs remains limited, particularly beyond the average treatment effect. In this paper, we derive the semiparametric efficiency bound for a class of propensity-score-weighted average treatment effects, which includes the average treatment effect, effects among treated and untreated populations, and the overlap effect, under two-phase sampling. In addition to straightforward weighting estimators based on the known sampling probabilities, we propose an enriched doubly robust estimator that attains the efficiency bound when all nuisance functions are consistently estimated. In particular, under outcome-dependent sampling, substantial efficiency gains can arise in some settings by appropriately incorporating phase-1 information. We further conduct extensive simulation studies, varying the choice of phase-1 variables and sampling schemes, to characterize when and to what extent leveraging phase-1 information leads to efficiency gains.

stat.ME

Identification and Estimation under Multiple Versions of Treatment: Mixture-of-Experts Approach

The Stable Unit Treatment Value Assumption (SUTVA) includes the condition that there are no multiple versions of treatment in causal inference. Though we could not control the implementation of treatment in observational studies, multiple versions may exist in the treatment. It has been pointed out that ignoring such multiple versions of treatment can lead to biased estimates of causal effects, but a causal inference framework that explicitly deals with the unbiased identification and estimation has not been fully developed yet. Thus, it is difficult to obtain a deeper understanding for mechanisms of the complex treatments. In this paper, we introduce the Mixture-of-Experts framework into causal inference to estimate causal contrasts between underlying versions of a treatment, even when the versions are not observed. Numerical experiments demonstrate the effectiveness of the proposed method.

stat.ME