Search arXivSearch

arXiv · 2112.02636

Optimal criteria and their asymptotic form for data selection in data-driven reduced-order modeling with Gaussian process regression

Abstract

We derive criteria for the selection of datapoints used for data-driven reduced-order modeling and other areas of supervised learning based on Gaussian process regression (GPR). While this is a well-studied area in the fields of active learning and optimal experimental design, most criteria in the literature are empirical. Here we introduce an optimality condition for the selection of a new input defined as the minimizer of the distance between the approximated output probability density function (pdf) of the reduced-order model and the exact one. Given that the exact pdf is unknown, we define the selection criterion as the supremum over the unit sphere of the native Hilbert space for the GPR. The resulting selection criterion, however, has a form that is difficult to compute. We combine results from GPR theory and asymptotic analysis to derive a computable form of the defined optimality criterion that is valid in the limit of small predictive variance. The derived asymptotic form of the selection criterion leads to convergence of the GPR model that guarantees a balanced distribution of data resources between probable and large-deviation outputs, resulting in an effective way for sampling towards data-driven reduced-order modeling.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Themistoklis P. Sapsis, Antoine Blanchard. 2021-12-05. Optimal criteria and their asymptotic form for data selection in data-driven reduced-order modeling with Gaussian process regression. https://doi.org/10.1098/rsta.2021.0197

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Equidistribution of saddle periodic points for Hénon-like maps

We prove that under a natural assumption on the dynamical degrees, the saddle periodic points of a Hénon-like map in any dimension equidistribute with respect to the equilibrium measure. Our work is a generalization of the results of Bedford-Lyubich-Smillie, Dujardin, and Dinh-Sibony along with improvements of their techniques. We also investigate some fine properties of Green currents associated with the map.

math.DS

On dissonance and orthogonal projections of self-conformal measures

Let $μ$ be a self-conformal measure on $\mathbb{R}^d$. We establish conditions for $μ$ under which $\dim(μ*ν) = \min\lbrace d,\dimμ+\dimν\rbrace$ holds when $ν$ is any Ahlfors-regular or self-conformal measure on $\mathbb{R}^d$. Our main result states the following sufficient condition: $μ$ is totally non-linear and not supported on a smooth hypersurface. We also establish sufficient (likely non-sharp) algebraic conditions for self-conformal measures which are not totally non-linear. In addition, we show that $\dim μ\circπ^{-1} = \min\{ k, \dim μ\}$ for every ortohogonal projection $π:\mathbb{R}^d\to\mathbb{R}^k$, $0<k<d$, when either $d=2$ and $μ$ is not self-similar and not supported on a line, or $d\geq 3$ and $μ$ is totally non-linear and not supported on a smooth hypersurface.

math.DS

Equation-Free Screening of Mittag-Leffler-Compatible Dynamics from Scalar Time Series via kNN Multi-Horizon Profiles

Fractional models provide a natural description of systems with memory, but a noninteger derivative should not be introduced solely because a time series is curved or slowly relaxing. We develop an equation-free preliminary screening framework that asks whether a scalar time series produces a multi-horizon k-nearest-neighbor (kNN) profile more compatible with Mittag-Leffler-type behavior than with selected conventional alternatives. In an ideal matched Caputo-relaxation benchmark, the complete generation-kNN-profile-model-comparison pipeline reproduces the expected Mittag-Leffler geometry and recovers the generating order to within approximately $10^{-3}$; this is interpreted as controlled calibration rather than as general fractional-order identification. Under 3% trajectory-specific observational noise, the held-out Mittag-Leffler preference is most consistent when the generating dynamics are well separated from the integer-order limit and becomes progressively less decisive as $α\rightarrow1$. The fitted order $α_{\mathrm{fit}}$, however, shows substantially larger realization-to-realization variability. Thus, relative model compatibility is more robust than single-realization order estimation in the present noisy benchmark. Noise-free nonfractional controls show a separate limitation of specificity: a stretched exponential can generate a strongly Mittag-Leffler-compatible profile, whereas inclusion of the generating rational/Hill family recovers that family and its parameters to numerical precision in the matched setting. A positive Mittag-Leffler-versus-exponential screen therefore does not uniquely establish fractional origin. A fractional chaotic system is treated only as an exploratory extension: the Mittag-Leffler growth family gives lower finite-window RMSE than exponential and logistic/saturating alternatives over the detected pre-transition interval.

math.DS