Search arXivSearch

arXiv · 2508.12206

The Identification Power of Combining Experimental and Observational Data for Distributional Treatment Effect Parameters

Abstract

This study investigates the identification power gained by combining experimental data, in which treatment is randomized, with observational data, in which treatment is self-selected, for distributional treatment effect (DTE) parameters. While experimental data identify average treatment effects, many DTE parameters, such as the distribution of individual treatment effects, are only partially identified. We examine whether and how combining these two data sources tightens the identified set for such parameters. For broad classes of DTE parameters, we derive nonparametric sharp bounds under the combined data and clarify the mechanism through which data combination improves identification relative to using experimental data alone. Our analysis highlights that self-selection in observational data is a key source of identification power. We establish necessary and sufficient conditions under which the combined data strictly shrink the identified set, and show that such gains arise generically unless selection-on-observables holds in the observational data. We also propose a linear programming approach to compute sharp bounds that can incorporate additional structural restrictions, such as positive dependence between potential outcomes and the generalized Roy selection model. An empirical application using data on negative campaign advertisements in the 2008 U.S. presidential election illustrates the practical relevance of the proposed approach.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shosei Sakaguchi. 2026-04-23. The Identification Power of Combining Experimental and Observational Data for Distributional Treatment Effect Parameters. https://arxiv.org/abs/2508.12206

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Poverty Targeting with Imperfect Information

How should antipoverty programs allocate transfers when household income is known only through noisy predictions? I formulate this as a statistical decision problem in which a policymaker chooses nonnegative transfers within a fixed budget to minimize squared deviations of post-transfer income from the poverty line. I show that the standard plug-in rule, which treats predictions as exact, is inadmissible. I then develop a nonparametric empirical Bayes allocation rule that replaces the estimated gaps with posterior mean poverty gaps. Its Bayes regret is bounded by the mean squared difference between its posterior mean gaps and the oracle's, so the budget and nonnegativity constraints do not slow its convergence to the oracle. The approach extends, with weaker guarantees, to a planner who cares only about how far households remain below the poverty line after transfers and to programs that pay a fixed set of benefit amounts. In simulations using household surveys from nine African countries, the allocations under the empirical Bayes rule reach about 1.8 times as many poor people as those under plug-in OLS with the same budget and achieve the same average poverty gap reduction with 6.7% less spending.

econ.EM

Estimating Treatment Effects in Panel Data Without Parallel Trends

This paper proposes a novel approach for estimating treatment effects in panel data settings, addressing key limitations of the standard difference-in-differences (DID) approach. The standard approach relies on the parallel trends assumption, implicitly requiring that unobservable factors correlated with treatment assignment be unidimensional, time-invariant, and affect untreated potential outcomes in an additively separable manner. This paper introduces a more flexible framework that allows for multidimensional unobservables and non-additive separability, and provides sufficient conditions for identifying the average treatment effect on the treated. An empirical application to job displacement reveals substantially smaller long-run earnings losses compared to the standard DID approach, demonstrating the framework's ability to account for unobserved heterogeneity that manifests as differential outcome trajectories between treated and control groups.

econ.EM

Single-Network Finite-Sample Inference in Strategic Network Formation Models

We develop a finite-sample valid inference procedure for strategic network formation models with endogenous network statistics, using only a single observed network. We impose no restrictions on network density, equilibrium selection, or the dependence structure induced by strategic interaction. Using a "bounding-by-c" technique, we obtain realization-wise sandwich inequalities whose middle term depends only on exogenous covariates and i.i.d. pairwise shocks. These inequalities deliver pathwise identifying restrictions and simulable finite-sample critical values in both parametric and semiparametric settings. Our proposed procedure is also computationally tractable: it does not involve solving, simulating, or enumerating equilibrium (sub)networks, and easily scales to networks with 10000 agents in simulations. In two empirical applications, we find statistical evidence for positive link interdependence at the 95% confidence level.

econ.EM