Search arXivSearch

arXiv · 1709.05289

Optimal approximation of piecewise smooth functions using deep ReLU neural networks

Abstract

We study the necessary and sufficient complexity of ReLU neural networks---in terms of depth and number of weights---which is required for approximating classifier functions in $L^2$. As a model class, we consider the set $\mathcal{E}^\beta (\mathbb R^d)$ of possibly discontinuous piecewise $C^\beta$ functions $f : [-1/2, 1/2]^d \to \mathbb R$, where the different smooth regions of $f$ are separated by $C^\beta$ hypersurfaces. For dimension $d \geq 2$, regularity $\beta > 0$, and accuracy $\varepsilon > 0$, we construct artificial neural networks with ReLU activation function that approximate functions from $\mathcal{E}^\beta(\mathbb R^d)$ up to $L^2$ error of $\varepsilon$. The constructed networks have a fixed number of layers, depending only on $d$ and $\beta$, and they have $O(\varepsilon^{-2(d-1)/\beta})$ many nonzero weights, which we prove to be optimal. In addition to the optimality in terms of the number of weights, we show that in order to achieve the optimal approximation rate, one needs ReLU networks of a certain depth. Precisely, for piecewise $C^\beta(\mathbb R^d)$ functions, this minimal depth is given---up to a multiplicative constant---by $\beta/d$. Up to a log factor, our constructed networks match this bound. This partly explains the benefits of depth for ReLU networks by showing that deep networks are necessary to achieve efficient approximation of (piecewise) smooth functions. Finally, we analyze approximation in high-dimensional spaces where the function $f$ to be approximated can be factorized into a smooth dimension reducing feature map $\tau$ and classifier function $g$---defined on a low-dimensional feature space---as $f = g \circ \tau$. We show that in this case the approximation rate depends only on the dimension of the feature space and not the input dimension.

Explore related subjects

Keep this discovery

BibTeXRIS

Philipp Petersen, Felix Voigtlaender. 2017-09-15. Optimal approximation of piecewise smooth functions using deep ReLU neural networks. https://arxiv.org/abs/1709.05289

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Typical dynamical properties of operators on $\ell_p$

We investigate the typical dynamical properties of hypercyclic operators in $\mathcal{L}_M(X)$, the set of all bounded linear operators on $X$ whose norms are at most $M$, when $X=\ell_p$, $1< p<\infty$. We show that, with respect to SOT$^*$, a typical operator $T\in \mathcal{L}_M(X)$ is weakly mixing, is weakly disjoint from a given hypercyclic operator $S$, is not topologically ergodic, and satisfies $(T,T^2,\dotsc,T^k)$ is disjoint hypercyclic for any $k\geq 2$. We also study the typical dynamical properties for the concrete family $\mathcal{M}=\{I+B_w\in \mathcal{L}(X)\colon w\in c_0(\mathbb{Z})\}$, endowed with the norm topology, where $B_w$ is a bilateral weighted backward shift.

math.FA

A bi-Lipschitz characterization of strong minimum-attainment for Lipschitz maps

We completely characterize the denseness of strongly minimum-attaining Lipschitz functions, a minimum analogue for strongly norm-attaining Lipschitz functions, in terms of bi-Lipschitz embeddings. More precisely, our main result shows that the set of strongly minimum-attaining Lipschitz functions defined on a complete metric space $M$ fails the denseness if and only if $M$ is bi-Lipschitz equivalent to a subset of $\mathbb{R}$ with positive Lebesgue measure, or equivalently, if $M$ admits a bi-Lipschitz embedding into $\mathbb{R}$ and $M$ has positive 1-dimensional Hausdorff measure. As a consequence, we provide an isometric characterization of the pure 1-unrectifiability of $M$ in terms of strongly minimum-attaining Lipschitz maps defined on bi-Lipschitz copies of closed subsets of $M$. Several counterexamples showing that the main result cannot be naturally extended to the vector-valued setting are also presented.

math.FA

On weak dominance of t-conorms over t-norms

The weak dominance of aggregation operators, particularly between triangular norms (t-norms) and triangular conorms (t-conorms), has attracted considerable attention in aggregation operator theory. While several characterizations have been obtained for Archimedean and continuous cases, a general criterion for continuous t-conorms over continuous t-norms remains to be fully clarified. In this paper, we provide a complete characterization of a continuous t-conorm weakly dominating a continuous t-norm. We first reduce the problem for ordinal sum operators to that for their single Archimedean components, and then express the weak dominance condition entirely in terms of the additive generators of these components. Our approach covers both strict and nilpotent cases uniformly, and recovers the known results for Archimedean operators as a special case.

math.FA