Search arXiv⌕ Search

arXiv · 0706.3923

Nonparametric estimation for dependent data with an application to panel time series

Abstract

In this paper we consider nonparametric estimation for dependent data, where the observations do not necessarily come from a linear process. We study density estimation and also discuss associated problems in nonparametric regression using the 2-mixing dependence measure. We compare the results under 2-mixing with those derived under the assumption that the process is linear. In the context of panel time series where one observes data from several individuals, it is often too strong to assume the joint linearity of processes. Instead the methods developed in this paper enable us to quantify the dependence through 2-mixing which allows for nonlinearity. We propose an estimator of the panel mean function and obtain its rate of convergence. We show that under certain conditions the rate of convergence can be improved by allowing the number of individuals in the panel to increase with time.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jan Johannes, Suhasini Subba Rao. 2007-06-27. Nonparametric estimation for dependent data with an application to panel time series. https://arxiv.org/abs/0706.3923

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Mallows-type model averaging: Non-asymptotic analysis and all-subset combination

Model averaging (MA) and ensembling play a crucial role in statistical and machine learning practice. When multiple candidate models are considered, MA techniques can be used to weight and combine them, often resulting in improved predictive accuracy and better estimation stability compared to model selection (MS) methods. In this paper, we address two important problems in combining least squares estimators from both theoretical and practical perspectives. We first establish several oracle inequalities for least squares MA via minimizing a Mallows' $C_p$ criterion under an arbitrary candidate model set. Compared to existing studies, these oracle inequalities yield faster excess risk and directly imply the asymptotic optimality of the resulting MA estimators under milder conditions. Moreover, we consider candidate model construction and investigate the problem of optimal all-subset combination for least squares estimators, which remains an important yet not fully understood topic in the existing literature. We show that there exists a fundamental limit to achieving the optimal all-subset MA risk. To attain this limit, we propose a novel Mallows-type MA procedure based on a dimension-adaptive $C_p$ criterion. The implicit ensembling effects of several MS procedures are also revealed and discussed. We conduct several numerical experiments to support our theoretical findings and demonstrate the effectiveness of the proposed Mallows-type MA estimator.

math.ST↗

Empirical Bernstein in smooth Banach spaces

Existing concentration bounds for bounded vector-valued random variables include extensions of the scalar Hoeffding and Bernstein inequalities. While the latter is typically tighter, it requires knowing a bound on the variance of the random variables. We derive a new vector-valued empirical Bernstein inequality, which makes use of an empirical estimator of the variance instead of the true variance. The bound holds in 2-smooth separable Banach spaces, which include finite dimensional Euclidean spaces and separable Hilbert spaces. The resulting confidence sets are instantiated for both the batch setting (where the sample size is fixed) and the sequential setting (where the sample size is a stopping time). The confidence set width asymptotically exactly matches that achieved by Bernstein in the leading term.

math.ST↗

Sharp Empirical Bernstein Bounds for the Variance of Bounded Random Variables

We develop novel empirical Bernstein inequalities for the variance of bounded random variables. Our inequalities hold under constant conditional variance and mean, without further assumptions like independence or identical distribution of the random variables, making them suitable for sequential decision making contexts. The results are instantiated for both the batch setting (where the sample size is fixed) and the sequential setting (where the sample size is a stopping time). Our bounds are asymptotically sharp: when the data are iid, our CI adpats optimally to both unknown mean $μ$ and unknown $\mathbb{V}[(X-μ)^2]$, meaning that the first order term of our CI exactly matches that of the oracle Bernstein inequality which knows those quantities. We compare our results to a widely used (non-sharp) concentration inequality for the variance based on self-bounding random variables, showing both the theoretical gains and improved empirical performance of our approach. We finally extend our methods to work in any separable Hilbert space.

math.ST↗