Search arXivSearch

arXiv · 2609.07888

Median-based Splitting Rules for Causal Trees and Forests

Abstract

Heavy-tailed and skewed outcomes are common in the randomized experiments and observational studies used to estimate heterogeneous treatment effects, yet the mean-squared-error criterion that guides splitting in honest causal trees is sensitive to the extreme values they generate. Building on the causal forest framework (Athey and Imbens, 2016; Wager and Athey, 2018), we introduce the Median Squared Deviation (MSD) criterion, which replaces the leafwise difference in means in the honest splitting objective with the Hodges--Lehmann location estimator while leaving honest leaf estimation and forest inference unchanged. Two further median-based rules, the Median Absolute Deviation (MAD) and the Least Median of Squares (LMS), serve as robust baselines. We evaluate the criteria in a simulation study covering precision, bias, and confidence interval coverage. MSD restricts its robustness to split selection and lowers the error of conditional average treatment effect estimates under heavy-tailed and skewed outcomes. Further, we re-visit two empirical applications: the first analyzes the electoral effects of a Mexican conditional cash transfer program on precinct-level observations, while the second application studies antiretroviral treatments in HIV-positive adults.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lennard Maßmann, Karolina Gliszczyńska-Schroeder. 2026-09-07. Median-based Splitting Rules for Causal Trees and Forests. https://arxiv.org/abs/2609.07888

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Poverty Targeting with Imperfect Information

How should antipoverty programs allocate transfers when household income is known only through noisy predictions? I formulate this as a statistical decision problem in which a policymaker chooses nonnegative transfers within a fixed budget to minimize squared deviations of post-transfer income from the poverty line. I show that the standard plug-in rule, which treats predictions as exact, is inadmissible. I then develop a nonparametric empirical Bayes allocation rule that replaces the estimated gaps with posterior mean poverty gaps. Its Bayes regret is bounded by the mean squared difference between its posterior mean gaps and the oracle's, so the budget and nonnegativity constraints do not slow its convergence to the oracle. The approach extends, with weaker guarantees, to a planner who cares only about how far households remain below the poverty line after transfers and to programs that pay a fixed set of benefit amounts. In simulations using household surveys from nine African countries, the allocations under the empirical Bayes rule reach about 1.8 times as many poor people as those under plug-in OLS with the same budget and achieve the same average poverty gap reduction with 6.7% less spending.

econ.EM

Estimating Treatment Effects in Panel Data Without Parallel Trends

This paper proposes a novel approach for estimating treatment effects in panel data settings, addressing key limitations of the standard difference-in-differences (DID) approach. The standard approach relies on the parallel trends assumption, implicitly requiring that unobservable factors correlated with treatment assignment be unidimensional, time-invariant, and affect untreated potential outcomes in an additively separable manner. This paper introduces a more flexible framework that allows for multidimensional unobservables and non-additive separability, and provides sufficient conditions for identifying the average treatment effect on the treated. An empirical application to job displacement reveals substantially smaller long-run earnings losses compared to the standard DID approach, demonstrating the framework's ability to account for unobserved heterogeneity that manifests as differential outcome trajectories between treated and control groups.

econ.EM

Single-Network Finite-Sample Inference in Strategic Network Formation Models

We develop a finite-sample valid inference procedure for strategic network formation models with endogenous network statistics, using only a single observed network. We impose no restrictions on network density, equilibrium selection, or the dependence structure induced by strategic interaction. Using a "bounding-by-c" technique, we obtain realization-wise sandwich inequalities whose middle term depends only on exogenous covariates and i.i.d. pairwise shocks. These inequalities deliver pathwise identifying restrictions and simulable finite-sample critical values in both parametric and semiparametric settings. Our proposed procedure is also computationally tractable: it does not involve solving, simulating, or enumerating equilibrium (sub)networks, and easily scales to networks with 10000 agents in simulations. In two empirical applications, we find statistical evidence for positive link interdependence at the 95% confidence level.

econ.EM