Search arXivSearch

arXiv · 2310.11962

Machine Learning for Staggered Difference-in-Differences and Dynamic Treatment Effect Heterogeneity

Abstract

We combine two recently proposed nonparametric difference-in-differences methods, extending them to enable the examination of treatment effect heterogeneity in the staggered adoption setting using machine learning. The proposed method, machine learning difference-in-differences (MLDID), allows for estimation of time-varying conditional average treatment effects on the treated, which can be used to conduct detailed inference on drivers of treatment effect heterogeneity. We perform simulations to evaluate the performance of MLDID and find that it accurately identifies the true predictors of treatment effect heterogeneity. We then use MLDID to evaluate the heterogeneous impacts of Brazil's Family Health Program on infant mortality, and find those in poverty and urban locations experienced the impact of the policy more quickly than other subgroups.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Julia Hatamyar, Noemi Kreif, Rudi Rocha, Martin Huber. 2023-10-18. Machine Learning for Staggered Difference-in-Differences and Dynamic Treatment Effect Heterogeneity. https://arxiv.org/abs/2310.11962

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Uniform Inference for Parameters Identified by Conditional Quantile Restrictions

Many structural and dynamic economic models imply that key parameters are identified by conditional quantile restrictions. Building on the exponential-weighting approach of Bierens (1990) and recent advances in penalized maximum statistics for conditional moment restrictions (Chen et al., 2025), we develop a unified inference framework for such parameters. We propose an adaptive $\ell_1$-penalized supremum statistic that transforms the conditional restriction into a continuum of unconditional moment conditions and aggregates evidence across quantile indices. The penalty regularizes the maximization over the weighting direction. Under the stated uniformity conditions, the known-parameter adaptive selector has no lower maximin local power than the unpenalized test and yields a strict maximin local-power gain whenever some positive candidate penalty has a strictly larger population maximin criterion than the zero penalty. We extend the theory to settings with pre-estimated nuisance parameters, characterizing the additional terms induced by the plug-in step in the limiting process. We derive an analytically corrected variance estimator that accounts for plug-in estimation uncertainty and establish the validity of a Gaussian multiplier bootstrap under the null and sequences of local alternatives. Monte Carlo simulations show that the proposed CvM-KS aggregation scheme has rejection rates relatively close to nominal size under pre-estimation in the linear design and approaches the nominal level in nonlinear designs as the sample size grows. The reported power comparisons between the adaptive and unpenalized procedures are design-specific, with higher empirical rejection frequencies for the adaptive procedure along some directions.

econ.EM

Recidivism Prediction, Peer Effect Estimation, and Prediction-Powered Inference with LLM Text Measures

We provide a new framework for estimating peer effects when outcomes are multivariate behavioral measures derived from written text using an LLM and the network formation is endogenous. We obtain LLM embeddings and zero shot classification of more than 200,000 written exchanges among residents of low-security correctional facilities. We find that LLM embeddings improve out-of-sample recidivism prediction by up to 30% over pre-entry covariates alone using LASSO and LoRA fine-tuning, showing that text representations capture meaningful signals. For peer effect estimation, we develop a novel instrumental variable estimator that accommodates multivariate outcomes, sparse networks, and multidimensional latent homophily. We show that this estimator is $\sqrt{N}$-consistent and asymptotically normal under sparsity conditions that relax dense-network assumptions prevalent in the peer effect literature. Limited human annotations are then combined with LLM zero-shot vectors in a new prediction-powered peer inference (PPPI) approach to obtain de-biased estimates and valid inference. Results reveal significant peer effects in the behavioral profiles.

econ.EM

Causal inference in two-sided randomization designs: factorial regression, two-way clustering, and covariate adjustment

We study randomized experiments involving two interacting populations, such as buyers and sellers in a marketplace. In the two-sided experiments we consider, we randomize the two populations separately and independently. For a pair consisting of one member from each population, the two assignments jointly determine one of four exposure conditions. Under a local interference assumption, we consider a broad class of linear estimands, including total, interaction, and buyer- and seller-side spillover effects. Our first main result establishes that researchers can estimate these effects using ordinary least squares and conduct asymptotically valid design-based inference using the conventional two-way cluster-robust variance estimator, clustered at the buyers' and sellers' levels. Our second main result develops a sharper variance estimator for a single linear estimand that better preserves dependence within the buyer and seller dimensions and is asymptotically less conservative than the two-way clustered estimator and existing alternatives. Our third main result establishes the theory for covariate adjustment and recommends a two-way analysis-of-variance-type covariate representation to ensure efficiency gains.

econ.EM