Search arXiv⌕ Search

arXiv · 1207.4421

Stochastic optimization and sparse statistical recovery: An optimal algorithm for high dimensions

Abstract

We develop and analyze stochastic optimization algorithms for problems in which the expected loss is strongly convex, and the optimum is (approximately) sparse. Previous approaches are able to exploit only one of these two structures, yielding an $\order(\pdim/T)$ convergence rate for strongly convex objectives in $\pdim$ dimensions, and an $\order(\sqrt{(\spindex \log \pdim)/T})$ convergence rate when the optimum is $\spindex$-sparse. Our algorithm is based on successively solving a series of $\ell_1$-regularized optimization problems using Nesterov's dual averaging algorithm. We establish that the error of our solution after $T$ iterations is at most $\order((\spindex \log\pdim)/T)$, with natural extensions to approximate sparsity. Our results apply to locally Lipschitz losses including the logistic, exponential, hinge and least-squares losses. By recourse to statistical minimax results, we show that our convergence rates are optimal up to multiplicative constant factors. The effectiveness of our approach is also confirmed in numerical simulations, in which we compare to several baselines on a least-squares regression problem.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Alekh Agarwal, Sahand Negahban, Martin J. Wainwright. 2012-07-18. Stochastic optimization and sparse statistical recovery: An optimal algorithm for high dimensions. https://arxiv.org/abs/1207.4421

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adaptive Subspace Modeling With Functional Tucker Decomposition

Tensors provide a structured representation for multidimensional data, yet discretization discards the underlying continuous structure when the data originate from continuous processes. We address this limitation by introducing a functional Tucker decomposition (FTD) that embeds a mode-wise continuity constraint directly into the factorization. The FTD models the continuous mode as a function in a reproducing kernel Hilbert space (RKHS), avoiding a prespecified basis while preserving the multilinear subspace structure of the Tucker model. We derive a reconstruction error bound for the continuous mode that quantifies the approximation quality when a subspace estimated on one domain is reused on another. This bound provides theoretical justification for subspace transfer, whose practical value we demonstrate on cross-domain classification tasks in hyperspectral imaging and multivariate time-series analysis.

stat.ML↗

Identifying Causal Effects Using a Single Proxy Variable

Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism that generates the proxy from the confounder. Under an assumption called Single Proxy Identifiability of Causal Effects or simply SPICE, we prove that this error mechanism is complete and causal effects are identifiable. We extend the proxy-based causal identifiability results by Kuroki and Pearl (2014); Pearl (2010) to multi-dimensional continuous settings, more flexible functional relationships and a broader class of distributions. Further, we develop a neural network based estimation framework, SPICE-Net, to estimate causal effects, which is applicable to both discrete and continuous treatments.

stat.ML↗

Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis

We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type parameterization the dual simplex matrix, reducing the trainable-state memory to $O(K^2)$, independent of $N$, and the cost per data pass to $O(NK^2)$. We prove a non-asymptotic sample-complexity bound of the polynomial-time benchmark order for a localized surrogate estimator; an oracle inequality for every global minimizer of the neural objective, with volume-inflation control and an explicit shrinkage bias; a conditional end-to-end error budget separating statistical, approximation, optimization, and enclosure-residual terms on an explicit envelope event; and two-point lower bounds: at any noise level $σ>0$ fixed independently of $N$, the $N^{-1/2}$ scaling is unimprovable in its $N$-exponent. Experiments with up to $N=10^8$ synthetic observations are consistent with the predicted accuracy and scaling, and feasibility on real scenes of $\sim 10^7$ pixels is demonstrated.

stat.ML↗