Search arXivSearch

arXiv · 2405.07860

Order-Explicit Linearization of High-Dimensional $U$-Statistics

Abstract

We give an order-explicit large deviation bound for the difference between a high-dimensional $U$-statistic and its Hájek projection. In particular, we show that any $U$-statistic of order $b$ on $n$ observations, with a $d$-dimensional kernel whose coordinates have $ψ_1$-Orlicz norm at most $ϕ$, has a maximum deviation from its Hájek projection of order $O_p(ϕb n^{-1}\log^2(dn))$. The proof relies on the development of novel order-explicit moment inequalities for higher-order Hoeffding components. We show that this rate is unimprovable, up to the polynomial factor on the logarithmic term. As corollaries, we obtain new Bernstein-type concentration and Gaussian approximation results for high-dimensional $U$-statistics. We apply these results to establish the consistency of a set of resampling-based simultaneous confidence intervals built around a class of nonparametric regression estimators constructed with subsampled kernels. This class encompasses several forms of random forest regression, including Generalized Random Forests.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David M. Ritzwoller, Vasilis Syrgkanis. 2026-07-13. Order-Explicit Linearization of High-Dimensional $U$-Statistics. https://arxiv.org/abs/2405.07860

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables

We propose a general approach to inference for a broad class of models that arise in the analysis of treatment effects with discrete-valued treatments and instruments and a general-valued outcome. In addition to instrument exogeneity, the main substantive assumption in our class of models rules out certain response types by assuming that they occur with probability zero. Here, the response type refers to the vector of potential outcomes and potential treatments, and we refer to a set of possible values for the response type as a generalized principal stratum. Through a series of examples, we show that this framework encompasses a wide variety of assumptions that have been considered in the previous literature. Our framework allows inference on any treatment effect parameter that can be expressed as the expectation of a function of the response type conditional on a generalized principal stratum. We develop methods for inference on such parameters under these assumptions, as well as methods for testing the validity of the assumptions themselves. A key result of our analysis is a characterization of the identified set for such parameters under these assumptions and the testable restrictions for the assumptions themselves in terms of existence of a nonnegative solution to linear systems of equations with a special structure. We propose methods for inference exploiting this special structure and recent results in Fang et al. (2023).

econ.EM

Raking for estimation and inference in panel models with nonignorable attrition and refreshment

In panel data subject to nonignorable attrition, auxiliary (refreshment) sampling may restore full identification under weak assumptions on the attrition process. Despite their generality, these identification strategies have seen limited empirical use, largely because the implied estimation procedure requires solving a functional minimization problem for the target density. When the distributions of the observed data are parametric, we show that this problem can be solved using the iterative proportional fitting (raking) algorithm, which converges rapidly even with continuous data. The resulting density estimator is then used as input into a parametric moment condition. We establish consistency and convergence rates for both the raking-based density estimator and the resulting moment estimator. We also derive a simple recursive procedure for estimating the asymptotic variance. Finally, we demonstrate the satisfactory performance of our estimator in simulations and provide an empirical illustration using data from the Understanding America Study panel.

econ.EM

Summary Indices in Treatment Effect Estimation

This paper studies the practice of combining multiple outcomes into a summary index to estimate a causal effect. For common estimators and index constructions, the estimate equals a weighted sum of the estimated effects on the components, with weights that are implicit and rarely reported. The paper derives the weights and shows that, for inverse-covariance-weighted indices, they can be negative and unrestricted in magnitude, so the index effect can have the opposite sign to every component effect. The paper proposes two procedures for valid inference on the index effect: a variance estimator that accounts for the data-dependent weights, and a shifted t-test that requires no such correction. Conventional t-tests of the null of no effect remain valid. Contrary to common claims, summary indices do not generally improve power. Three published studies illustrate the results.

econ.EM