Search arXivSearch

arXiv · 1910.08273

Large Dimensional Latent Factor Modeling with Missing Observations and Applications to Causal Inference

Abstract

This paper develops the inferential theory for latent factor models estimated from large dimensional panel data with missing observations. We propose an easy-to-use all-purpose estimator for a latent factor model by applying principal component analysis to an adjusted covariance matrix estimated from partially observed panel data. We derive the asymptotic distribution for the estimated factors, loadings and the imputed values under an approximate factor model and general missing patterns. The key application is to estimate counterfactual outcomes in causal inference from panel data. The unobserved control group is modeled as missing values, which are inferred from the latent factor model. The inferential theory for the imputed values allows us to test for individual treatment effects at any time under general adoption patterns where the units can be affected by unobserved factors.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ruoxuan Xiong, Markus Pelger. 2022-01-10. Large Dimensional Latent Factor Modeling with Missing Observations and Applications to Causal Inference. https://arxiv.org/abs/1910.08273

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Changes-in-Changes for Ordered Choice Models with Underreporting

We develop a Difference-in-Differences framework for discrete, ordered outcomes subject to underreporting. Such outcomes commonly arise in self-reported surveys on socially undesirable or stigmatized behaviors, where respondents may conceal their true behavior. For a discrete Changes-in-Changes model that is shown to admit an equivalent threshold-crossing representation, we derive nonparametric bounds for the counterfactual and factual outcome distributions as well as for the associated quantile treatment effects when outcomes are underreported. These bounds are shown to be sharp uniformly across outcome levels under additional support conditions, and we propose suitable estimation and bootstrap inference procedures. In an extension, we also consider a semiparametric underreporting model that allows to point identify and estimate distributional treatment effects. As an application, we investigate the impact of recreational marijuana legalization on the consumption behavior of 8th-grade students in several U.S. states.

econ.EM

Triple Difference Designs with Heterogeneous Treatment Effects

Triple difference designs have become increasingly popular in empirical economics. The advantage of a triple difference design is that, within a treatment group, it allows another subgroup of the population -- potentially less impacted by the treatment -- to serve as a comparison for the subgroup of interest. While literature on difference-in-differences has discussed heterogeneity in treatment effects between treated and control groups or over time, relatively little attention has been given to triple difference designs and the implications of heterogeneity in treatment effects in this setting. In this paper, I show that the parameter identified under common triple difference assumptions does not allow for causal interpretation of differences between subgroups when subgroups may differ in their underlying (unobserved) treatment effects. I propose a new parameter of interest, the controlled difference in average treatment effects on the treated, which allows for causal comparisons between subgroups. I then propose identification assumptions and doubly-robust estimators for this parameter. I use a simulation study to highlight the desirable finite-sample properties of these estimators, as well as to show the difference between the two parameters. An empirical application shows the importance of considering treatment effect heterogeneity in practical applications.

econ.EM

Generalized Covariance Estimator under Misspecification

This paper investigates the properties of the Generalized Covariance (GCov) estimator under misspecification with application to processes with local explosive patterns, such as causal-noncausal processes. We show that GCov is consistent and has an asymptotically Normal distribution under misspecification. Then, we construct GCov-based Wald-type and score-type tests to test one specification against the other, all of which follow a $χ^2$ distribution. We validate the finite-sample performance of the proposed estimators and tests in the context of causal-noncausal models. Finally, we provide applications of the noncausal model to the final energy demand commodity index.

econ.EM