Search arXiv⌕ Search

arXiv · 2009.04324

Overcoming the curse of dimensionality with Laplacian regularization in semi-supervised learning

Abstract

As annotations of data can be scarce in large-scale practical problems, leveraging unlabelled examples is one of the most important aspects of machine learning. This is the aim of semi-supervised learning. To benefit from the access to unlabelled data, it is natural to diffuse smoothly knowledge of labelled data to unlabelled one. This induces to the use of Laplacian regularization. Yet, current implementations of Laplacian regularization suffer from several drawbacks, notably the well-known curse of dimensionality. In this paper, we provide a statistical analysis to overcome those issues, and unveil a large body of spectral filtering methods that exhibit desirable behaviors. They are implemented through (reproducing) kernel methods, for which we provide realistic computational guidelines in order to make our method usable with large amounts of data.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vivien Cabannes, Loucas Pillaud-Vivien, Francis Bach, Alessandro Rudi. 2021-11-29. Overcoming the curse of dimensionality with Laplacian regularization in semi-supervised learning. https://arxiv.org/abs/2009.04324

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adaptive Subspace Modeling With Functional Tucker Decomposition

Tensors provide a structured representation for multidimensional data, yet discretization discards the underlying continuous structure when the data originate from continuous processes. We address this limitation by introducing a functional Tucker decomposition (FTD) that embeds a mode-wise continuity constraint directly into the factorization. The FTD models the continuous mode as a function in a reproducing kernel Hilbert space (RKHS), avoiding a prespecified basis while preserving the multilinear subspace structure of the Tucker model. We derive a reconstruction error bound for the continuous mode that quantifies the approximation quality when a subspace estimated on one domain is reused on another. This bound provides theoretical justification for subspace transfer, whose practical value we demonstrate on cross-domain classification tasks in hyperspectral imaging and multivariate time-series analysis.

stat.ML↗

Identifying Causal Effects Using a Single Proxy Variable

Unobserved confounding is a key challenge when estimating causal effects from a treatment on an outcome. In this work, we assume that we observe a single, potentially multi-dimensional proxy variable of the unobserved confounder and that we know the mechanism that generates the proxy from the confounder. Under an assumption called Single Proxy Identifiability of Causal Effects or simply SPICE, we prove that this error mechanism is complete and causal effects are identifiable. We extend the proxy-based causal identifiability results by Kuroki and Pearl (2014); Pearl (2010) to multi-dimensional continuous settings, more flexible functional relationships and a broader class of distributions. Further, we develop a neural network based estimation framework, SPICE-Net, to estimate causal effects, which is applicable to both discrete and continuous treatments.

stat.ML↗

Scalable Minimum-Volume Simplex Estimation with Non-asymptotic Analysis

We study the estimation of a $K$-dimensional simplex from $N$ i.i.d.\ points sampled uniformly from its interior; the observations are convex combinations of $K+1$ unknown prototypes. Existing polynomial-time estimators need cubic per-sample work or $O(NK)$ storage and are impractical at $N\sim 10^6$--$10^8$. We propose DeepMVSA, which re-expresses the minimum-volume principle in neural implicit form: a lightweight coordinate network generates the mixing weights and a triangular LU-type parameterization the dual simplex matrix, reducing the trainable-state memory to $O(K^2)$, independent of $N$, and the cost per data pass to $O(NK^2)$. We prove a non-asymptotic sample-complexity bound of the polynomial-time benchmark order for a localized surrogate estimator; an oracle inequality for every global minimizer of the neural objective, with volume-inflation control and an explicit shrinkage bias; a conditional end-to-end error budget separating statistical, approximation, optimization, and enclosure-residual terms on an explicit envelope event; and two-point lower bounds: at any noise level $σ>0$ fixed independently of $N$, the $N^{-1/2}$ scaling is unimprovable in its $N$-exponent. Experiments with up to $N=10^8$ synthetic observations are consistent with the predicted accuracy and scaling, and feasibility on real scenes of $\sim 10^7$ pixels is demonstrated.

stat.ML↗