Search arXiv⌕ Search

arXiv · 2610.01464

Inference after data-driven control-unit selection in difference-in-differences with estimated covariance

Abstract

In difference-in-differences (DiD), researchers may use pre-treatment trends to select a control group for which the parallel-trends assumption appears plausible, with the aim of estimating the average treatment effect on the treated (ATT). Our earlier paper,Nakano and Hoshino (2016), and the present paper jointly provide the first selective-inference approach to the ATT that explicitly accounts for this control selection. We generalize our exact Gaussian procedure with known covariance to allow the covariance matrix to be estimated from the same individual-level data used for control selection and DiD estimation. We use this estimate to compute the variance, conditioning direction, residual, and truncation set. With fixed numbers of regions and periods, we establish uniform conditional coverage for selection events with probabilities bounded away from zero, and marginal coverage of the selected target without that restriction. We allow unequal regional sample sizes, heterogeneous covariances, ties in population fit, and regional sample shares that converge to zero. We establish asymptotic equivalence between the plug-in and known-covariance interval endpoints and derive rates for interval length. For staggered adoption, the control pools may differ across cohorts and periods, controls may be not yet treated, observations may be reused, and treatment effects may be heterogeneous. We also construct inference conditional on unions of selection paths that leave the reported parameter unchanged, together with simultaneous confidence bands for finitely many event-time effects. Under parallel trends and the other identifying conditions, the coverage results apply to the ATT. We give sufficient sampling conditions for individual panels and independent repeated cross-sections.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ryoya Nakano, Takahiro Hoshino. 2026-10-01. Inference after data-driven control-unit selection in difference-in-differences with estimated covariance. https://arxiv.org/abs/2610.01464

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Nowcasting using regression on signatures

We introduce a new method of nowcasting using regression on path signatures. Path signatures capture the geometric properties of sequential data. Because signatures embed observations in continuous time, they naturally handle mixed frequencies and missing data. We prove theoretically, and demonstrate with simulations, that regression on signatures both subsumes the linear Kalman filter and has desirable consistency properties. Nowcasting with signatures is more robust to disruptions in data series than previous methods, making it useful in stressed times (for example, during COVID-19). This approach is performant in nowcasting US GDP growth, and in nowcasting UK unemployment.

econ.EM↗

Externally Valid Selection of Experimental Sites via the k-Median Problem

We present a decision-theoretic justification for viewing the question of how to best choose where to experiment in order to optimize external validity as a $k$-median problem, a popular problem in computer science and operations research. In particular, when treatment effect heterogeneity across experimental and policy-relevant sites is substantial (in a sense we make precise), we present conditions under which minimizing the worst-case, welfare-based regret among all nonrandom schemes that select $k$ sites to experiment is equivalent to solving a $k$-median problem. The connection costs in the relevant $k$-median problem are given by ex-ante bounds on worst-case voltage effects between sites, and minimizing the sum of worst-case voltage effects can be cast as a linear integer program. Two empirical applications illustrate the theoretical and computational benefits of the suggested procedure.

econ.EM↗

ACT, WAIT, or EXPERIMENT: A Causal Governance Framework for Retail Price Optimization Under Abstentions

This paper presents a causal decision-making framework for estimating price elasticity in retail channels, a process typically confounded by promotions, competitor movements, and market frictions. Rather than forcing a calculation when data is ambiguous, the system introduces decision abstention (\textsc{wait}) as an active diagnostic tool rather than an estimation failure. Combining Double Machine Learning and conformal prediction, the tool evaluates whether reliable conditions exist to adjust prices or if pausing the decision is preferable. When the system abstains, it exhaustively classifies the reason for the pause, identifying which products require designed pricing experiments or whether aggregating data to the brand level restores usable estimates. Tested on controlled synthetic data, the model shows that this operational discipline drastically reduces estimation error (lowering RMSE from 0.571 to 0.159) and offers a practical, secure alternative to blind estimation in thin-data retail environments.

econ.EM↗