Search arXiv⌕ Search

arXiv · 2009.03151

Doubly Robust Semiparametric Difference-in-Differences Estimators with High-Dimensional Data

Abstract

This paper proposes a doubly robust two-stage semiparametric difference-in-difference estimator for estimating heterogeneous treatment effects with high-dimensional data. Our new estimator is robust to model miss-specifications and allows for, but does not require, many more regressors than observations. The first stage allows a general set of machine learning methods to be used to estimate the propensity score. In the second stage, we derive the rates of convergence for both the parametric parameter and the unknown function under a partially linear specification for the outcome equation. We also provide bias correction procedures to allow for valid inference for the heterogeneous treatment effects. We evaluate the finite sample performance with extensive simulation studies. Additionally, a real data analysis on the effect of Fair Minimum Wage Act on the unemployment rate is performed as an illustration of our method. An R package for implementing the proposed method is available on Github.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yang Ning, Sida Peng, Jing Tao. 2020-09-07. Doubly Robust Semiparametric Difference-in-Differences Estimators with High-Dimensional Data. https://arxiv.org/abs/2009.03151

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Nonparametric Identification of First-Price Auction with Unobserved Competition: A Density Discontinuity Framework

We consider nonparametric identification of independent private value first-price auction models, in which the analyst only observes winning bids. Our benchmark model assumes an exogenous number of bidders N. We show that, if the bidders observe N, the resulting discontinuities in the winning bid density can be used to identify the distribution of N. The private value distribution can be nonparametrically identified in a second step. This extends, under testable identification conditions, to the case where N is a number of potential buyers, who bid with some unknown probability. Identification also holds in presence of additive unobserved heterogeneity drawn from some parametric distributions.

econ.EM↗

Welfare at Risk: Distributionally Robust Bounds for Policy Evaluation

This paper develops a framework to study how the welfare effects of policy interventions are distributed across individuals when those effects are not directly observed. We bound the superquantile of individual welfare changes by the superquantile of conditional average welfare changes, which is typically identified, revealing how gains and losses are spread across the population. We also bound the share of the population whose welfare loss from the policy exceeds any given threshold --- a distribution-free counterpart to average cost--benefit analysis, closer to a poverty-rate measure of policy harm than to a mean welfare effect. We further extend both bounds to allow for ambiguity in the analyst's knowledge of the distribution of observable characteristics: a non-robust bound estimated on one population can be systematically overoptimistic when applied to a different target population, while the robust version, with its ambiguity radius chosen commensurate with the shift, restores a valid guarantee. We illustrate the method in generalized Roy models with subjective participation costs, bounding not only the distribution of welfare changes in the overall population but also, in a result new to this literature, the entire distribution of welfare (and of participation costs) among those who select into treatment --- the treatment-on-the-treated distribution, not merely its mean. We connect these results to the multi-outcome IV framework of \cite{Heckman_Urzua_Vytlacil_2006_REStat}, showing our robust bounds are immune to the instrument-sensitivity problem they document

econ.EM↗

Corrected Forecast Combinations

This paper proposes corrected forecast combinations when the original combined forecast errors are serially dependent. Motivated by the classic Bates and Granger (1969) example, we show that combined forecast errors can be strongly autocorrelated and that a simple correction, which adds a fraction of the previous combined error to the next-period combined forecast, can deliver sizable improvements in forecast accuracy, often exceeding the original gains from combining. We formalize the approach within the conditional-risk framework of Gibbs and Vasnev (2024), in which the combined error decomposes into a predictable component (measurable at the forecast origin) and an innovation. We then link this correction to efficient estimation of combination weights under time-series dependence via GLS, allowing joint estimation of weights and an error-covariance structure. Using the U.S. Survey of Professional Forecasters for major macroeconomic indices across various subsamples (including pre/post-2000, GFC, and COVID), we find that a parsimonious correction of the mean forecast with a coefficient around 0.5 is a robust starting point and often yields material improvements in forecast accuracy. The benefits of correction diminish rapidly as the information becomes stale, largely disappearing at the four-quarter horizon or when the correction relies on an older lagged forecast error. For optimal-weight forecasts, the correction substantially mitigates the forecast combination puzzle by turning poorly performing out-of-sample optimal-weight combinations into competitive forecasts.

econ.EM↗