Search arXivSearch

arXiv · 2204.05175

Partially Linear Models under Data Combination

Abstract

We study partially linear models when the outcome of interest and some of the covariates are observed in two different datasets that cannot be linked. This type of data combination problem arises very frequently in empirical microeconomics. Using recent tools from optimal transport theory, we derive a constructive characterization of the sharp identified set. We then build on this result and develop a novel inference method that exploits the specific geometric properties of the identified set. Our method exhibits good performances in finite samples, while remaining very tractable. We apply our approach to study intergenerational income mobility over the period 1850-1930 in the United States. Our method allows us to relax the exclusion restrictions used in earlier work, while delivering confidence regions that are informative.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Xavier D'Haultfœuille, Christophe Gaillac, Arnaud Maurel. 2023-08-22. Partially Linear Models under Data Combination. https://arxiv.org/abs/2204.05175

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Poverty Targeting with Imperfect Information

How should antipoverty programs allocate transfers when household income is known only through noisy predictions? I formulate this as a statistical decision problem in which a policymaker chooses nonnegative transfers within a fixed budget to minimize squared deviations of post-transfer income from the poverty line. I show that the standard plug-in rule, which treats predictions as exact, is inadmissible. I then develop a nonparametric empirical Bayes allocation rule that replaces the estimated gaps with posterior mean poverty gaps. Its Bayes regret is bounded by the mean squared difference between its posterior mean gaps and the oracle's, so the budget and nonnegativity constraints do not slow its convergence to the oracle. The approach extends, with weaker guarantees, to a planner who cares only about how far households remain below the poverty line after transfers and to programs that pay a fixed set of benefit amounts. In simulations using household surveys from nine African countries, the allocations under the empirical Bayes rule reach about 1.8 times as many poor people as those under plug-in OLS with the same budget and achieve the same average poverty gap reduction with 6.7% less spending.

econ.EM

Estimating Treatment Effects in Panel Data Without Parallel Trends

This paper proposes a novel approach for estimating treatment effects in panel data settings, addressing key limitations of the standard difference-in-differences (DID) approach. The standard approach relies on the parallel trends assumption, implicitly requiring that unobservable factors correlated with treatment assignment be unidimensional, time-invariant, and affect untreated potential outcomes in an additively separable manner. This paper introduces a more flexible framework that allows for multidimensional unobservables and non-additive separability, and provides sufficient conditions for identifying the average treatment effect on the treated. An empirical application to job displacement reveals substantially smaller long-run earnings losses compared to the standard DID approach, demonstrating the framework's ability to account for unobserved heterogeneity that manifests as differential outcome trajectories between treated and control groups.

econ.EM

Single-Network Finite-Sample Inference in Strategic Network Formation Models

We develop a finite-sample valid inference procedure for strategic network formation models with endogenous network statistics, using only a single observed network. We impose no restrictions on network density, equilibrium selection, or the dependence structure induced by strategic interaction. Using a "bounding-by-c" technique, we obtain realization-wise sandwich inequalities whose middle term depends only on exogenous covariates and i.i.d. pairwise shocks. These inequalities deliver pathwise identifying restrictions and simulable finite-sample critical values in both parametric and semiparametric settings. Our proposed procedure is also computationally tractable: it does not involve solving, simulating, or enumerating equilibrium (sub)networks, and easily scales to networks with 10000 agents in simulations. In two empirical applications, we find statistical evidence for positive link interdependence at the 95% confidence level.

econ.EM