Search arXivSearch

arXiv subjects

Pascal Pernot

Publications and source records attributed to Pascal Pernot.

At least 19 recordsLinked to original sources

On the good reliability of an interval-based metric to validate prediction uncertainty for machine learning regression tasks

This short study presents an opportunistic approach to a (more) reliable validation method for prediction uncertainty average calibration. Considering that variance-based calibration metrics (ZMS, NLL, RCE...) are quite sensitive to the presence of heavy tails in the uncertainty and error distributions, a shift is proposed to an interval-based metric, the Prediction Interval Coverage Probability (PICP). It is shown on a large ensemble of molecular properties datasets that (1) sets of z-scores are well represented by Student's-$t(\nu)$ distributions, $\nu$ being the number of degrees of freedom; (2) accurate estimation of 95 $\%$ prediction intervals can be obtained by the simple $2\sigma$ rule for $\nu>3$; and (3) the resulting PICPs are more quickly and reliably tested than variance-based calibration metrics. Overall, this method enables to test 20 $\%$ more datasets than ZMS testing. Conditional calibration is also assessed using the PICP approach.

stat.ML

Validation of ML-UQ calibration statistics using simulated reference values: a sensitivity analysis

Some popular Machine Learning Uncertainty Quantification (ML-UQ) calibration statistics do not have predefined reference values and are mostly used in comparative studies. In consequence, calibration is almost never validated and the diagnostic is left to the appreciation of the reader. Simulated reference values, based on synthetic calibrated datasets derived from actual uncertainties, have been proposed to palliate this problem. As the generative probability distribution for the simulation of synthetic errors is often not constrained, the sensitivity of simulated reference values to the choice of generative distribution might be problematic, shedding a doubt on the calibration diagnostic. This study explores various facets of this problem, and shows that some statistics are excessively sensitive to the choice of generative distribution to be used for validation when the generative distribution is unknown. This is the case, for instance, of the correlation coefficient between absolute errors and uncertainties (CC) and of the expected normalized calibration error (ENCE). A robust validation workflow to deal with simulated reference values is proposed.

stat.ML

Negative impact of heavy-tailed uncertainty and error distributions on the reliability of calibration statistics for machine learning regression tasks

Average calibration of the (variance-based) prediction uncertainties of machine learning regression tasks can be tested in two ways: one is to estimate the calibration error (CE) as the difference between the mean absolute error (MSE) and the mean variance (MV); the alternative is to compare the mean squared z-scores (ZMS) to 1. The problem is that both approaches might lead to different conclusions, as illustrated in this study for an ensemble of datasets from the recent machine learning uncertainty quantification (ML-UQ) literature. It is shown that the estimation of MV, MSE and their confidence intervals becomes unreliable for heavy-tailed uncertainty and error distributions, which seems to be a frequent feature of ML-UQ datasets. By contrast, the ZMS statistic is less sensitive and offers the most reliable approach in this context, still acknowledging that datasets with heavy-tailed z-scores distributions should be considered with great care. Unfortunately, the same problem is expected to affect also conditional calibrations statistics, such as the popular ENCE, and very likely post-hoc calibration methods based on similar statistics. Several solutions to circumvent the outlined problems are proposed.

stat.ML

Can bin-wise scaling improve consistency and adaptivity of prediction uncertainty for machine learning regression ?

Binwise Variance Scaling (BVS) has recently been proposed as a post hoc recalibration method for prediction uncertainties of machine learning regression problems that is able of more efficient corrections than uniform variance (or temperature) scaling. The original version of BVS uses uncertainty-based binning, which is aimed to improve calibration conditionally on uncertainty, i.e. consistency. I explore here several adaptations of BVS, in particular with alternative loss functions and a binning scheme based on an input-feature (X) in order to improve adaptivity, i.e. calibration conditional on X. The performances of BVS and its proposed variants are tested on a benchmark dataset for the prediction of atomization energies and compared to the results of isotonic regression.

stat.ML

Calibration in Machine Learning Uncertainty Quantification: beyond consistency to target adaptivity

Reliable uncertainty quantification (UQ) in machine learning (ML) regression tasks is becoming the focus of many studies in materials and chemical science. It is now well understood that average calibration is insufficient, and most studies implement additional methods testing the conditional calibration with respect to uncertainty, i.e. consistency. Consistency is assessed mostly by so-called reliability diagrams. There exists however another way beyond average calibration, which is conditional calibration with respect to input features, i.e. adaptivity. In practice, adaptivity is the main concern of the final users of a ML-UQ method, seeking for the reliability of predictions and uncertainties for any point in features space. This article aims to show that consistency and adaptivity are complementary validation targets, and that a good consistency does not imply a good adaptivity. Adapted validation methods are proposed and illustrated on a representative example.

stat.ML

Stratification of uncertainties recalibrated by isotonic regression and its impact on calibration error statistics

Abstract Post hoc recalibration of prediction uncertainties of machine learning regression problems by isotonic regression might present a problem for bin-based calibration error statistics (e.g. ENCE). Isotonic regression often produces stratified uncertainties, i.e. subsets of uncertainties with identical numerical values. Partitioning of the resulting data into equal-sized bins introduces an aleatoric component to the estimation of bin-based calibration statistics. The partitioning of stratified data into bins depends on the order of the data, which is typically an uncontrolled property of calibration test/validation sets. The tie-braking method of the ordering algorithm used for binning might also introduce an aleatoric component. I show on an example how this might significantly affect the calibration diagnostics.

stat.ME

Properties of the ENCE and other MAD-based calibration metrics

The Expected Normalized Calibration Error (ENCE) is a popular calibration statistic used in Machine Learning to assess the quality of prediction uncertainties for regression problems. Estimation of the ENCE is based on the binning of calibration data. In this short note, I illustrate an annoying property of the ENCE, i.e. its proportionality to the square root of the number of bins for well calibrated or nearly calibrated datasets. A similar behavior affects the calibration error based on the variance of z-scores (ZVE), and in both cases this property is a consequence of the use of a Mean Absolute Deviation (MAD) statistic to estimate calibration errors. Hence, the question arises of which number of bins to choose for a reliable estimation of calibration error statistics. A solution is proposed to infer ENCE and ZVE values that do not depend on the number of bins for datasets assumed to be calibrated, providing simultaneously a statistical calibration test. It is also shown that the ZVE is less sensitive than the ENCE to outstanding errors or uncertainties.

cs.LG

Validation of uncertainty quantification metrics: a primer based on the consistency and adaptivity concepts

The practice of uncertainty quantification (UQ) validation, notably in machine learning for the physico-chemical sciences, rests on several graphical methods (scattering plots, calibration curves, reliability diagrams and confidence curves) which explore complementary aspects of calibration, without covering all the desirable ones. For instance, none of these methods deals with the reliability of UQ metrics across the range of input features (adaptivity). Based on the complementary concepts of consistency and adaptivity, the toolbox of common validation methods for variance- and intervals- based UQ metrics is revisited with the aim to provide a better grasp on their capabilities. This study is conceived as an introduction to UQ validation, and all methods are derived from a few basic rules. The methods are illustrated and tested on synthetic datasets and representative examples extracted from the recent physico-chemical machine learning UQ literature.

physics.chem-ph

Confidence curves for UQ validation: probabilistic reference vs. oracle

Confidence curves are used in uncertainty validation to assess how large uncertainties ($u_{E}$) are associated with large errors ($E$). An oracle curve is commonly used as reference to estimate the quality of the tested datasets. The oracle is a perfect, deterministic, error predictor, such as $|E|=\pm u_{E}$, which corresponds to a very unlikely error distribution in a probabilistic framework and is unable unable to inform us on the calibration of $u_{E}$. I propose here to replace the oracle by a probabilistic reference curve, deriving from the more realistic scenario where errors should be random draws from a distribution with standard deviation $u_{E}$. The probabilistic curve and its confidence interval enable a direct test of the quality of a confidence curve. Paired with the probabilistic reference, a confidence curve can be used to check the calibration and tightness of prediction uncertainties.

physics.data-an

Prediction uncertainty validation for computational chemists

Validation of prediction uncertainty (PU) is becoming an essential task for modern computational chemistry. Designed to quantify the reliability of predictions in meteorology, the calibration-sharpness (CS) framework is now widely used to optimize and validate uncertainty-aware machine learning (ML) methods. However, its application is not limited to ML and it can serve as a principled framework for any PU validation. The present article is intended as a step-by-step introduction to the concepts and techniques of PU validation in the CS framework, adapted to the specifics of computational chemistry. The presented methods range from elementary graphical checks to more sophisticated ones based on local calibration statistics. The concept of tightness, is introduced. The methods are illustrated on synthetic datasets and applied to uncertainty quantification data extracted from the computational chemistry literature.

physics.chem-ph

The long road to calibrated prediction uncertainty in computational chemistry

Uncertainty quantification (UQ) in computational chemistry (CC) is still in its infancy. Very few CC methods are designed to provide a confidence level on their predictions, and most users still rely improperly on the mean absolute error as an accuracy metric. The development of reliable uncertainty quantification methods is essential, notably for computational chemistry to be used confidently in industrial processes. A review of the CC-UQ literature shows that there is no common standard procedure to report nor validate prediction uncertainty. I consider here analysis tools using concepts (calibration and sharpness) developed in meteorology and machine learning for the validation of probabilistic forecasters. These tools are adapted to CC-UQ and applied to datasets of prediction uncertainties provided by composite methods, Bayesian Ensembles methods, machine learning and a posteriori statistical methods.

physics.chem-ph

Objective assessment of corneal transparency in the clinical setting with standard SD-OCT devices

PURPOSE: To develop an automated algorithm allowing extraction of quantitative corneal transparency parameters from clinical spectral-domain OCT images. To establish a representative dataset of normative transparency values from healthy corneas. METHODS: SD-OCT images of 83 normal corneas (ages 22-50 years) from a standard clinical device (RTVue-XR Avanti, Optovue Inc.) were processed. A pre-processing procedure is applied first, including a derivative approach and a PCA-based correction mask, to eliminate common central artifacts (i.e., apex-centered column saturation artifact and posterior stromal artifact) and enable standardized analysis. The mean intensity stromal-depth profile is then extracted over a 6-mm-wide corneal area and analyzed according to our previously developed method deriving quantitative transparency parameters related to the physics of light propagation in tissues, notably tissular heterogeneity (Birge ratio; $B_r$), followed by the photon mean-free path ($l_s$) in homogeneous tissues (i.e., $B_r \sim 1$). RESULTS: After confirming stromal homogeneity ($B_r < 10$, IDR: 1.9-5.1), we measured a median $l_s$ of 570 $\mu$m (IDR: 270-2400 $\mu$m). Considering corneal thicknesses, this may be translated into a median fraction of transmitted (coherent) light $T_{coh(stroma)}$ of 51$\%$ (IDR: 22-83$\%$). No statistically significant correlation between transparency and age or thickness was found. CONCLUSIONS: Our algorithm provides robust and quantitative measurement of corneal transparency from standard clinical SD-OCT images. It yields lower transparency values than previously reported, which may be attributed to our method being exclusively sensitive to spatially coherent light. Excluding images with central artifacts wider than 300 $\mu$m also raises our median $T_{coh(stroma)}$ to 70$\%$ (IDR: 34-87$\%$).

physics.med-ph

Using the Gini coefficient to characterize the shape of computational chemistry error distributions

The distribution of errors is a central object in the assesment and benchmarking of computational chemistry methods. The popular and often blind use of the mean unsigned error as a benchmarking statistic leads to ignore distributions features that impact the reliability of the tested methods. We explore how the Gini coefficient offers a global representation of the errors distribution, but, except for extreme values, does not enable an unambiguous diagnostic. We propose to relieve the ambiguity by applying the Gini coefficient to mode-centered error distributions. This version can usefully complement benchmarking statistics and alert on error sets with potentially problematic shapes.

physics.chem-ph

Ions in the Thermosphere of Exoplanets: Observable Constraints Revealed by Innovative Laboratory Experiments

With the upcoming launch of space telescopes dedicated to the study of exoplanets, the \textit{Atmospheric Remote-Sensing Infrared Exoplanet Large-survey} (ARIEL) and the \textit{James Webb Space Telescope} (JWST), a new era is opening in the exoplanetary atmospheric explorations. However, especially in relatively cold planets around later-type stars, photochemical hazes and clouds may mask the composition of the lower part of the atmosphere, making it difficult to detect any chemical species in the troposphere or to understand whether there is a surface or not. This issue is particularly exacerbated if the goal is to study the habitability of said exoplanets and to search for biosignatures.\par This work combines innovative laboratory experiments, chemical modeling and simulated observations at ARIEL and JWST resolutions. We focus on the signatures of molecular ions that can be found in upper atmospheres above cloud decks. Our results suggest that H$_3^+$ along with H$_3$O$^+$ could be detected in the observational spectra of sub-Neptunes based on realistic mixing ratio assumption. This new parametric set may help to distinguish super-Earths with a thin atmosphere from H$_2$-dominated sub-Neptunes, to address the critical question whether a low-gravity planet around a low-mass active star is able to retain its volatile components. These ions may also constitute potential tracers to certain molecules of interest such as H$_2$O or O$_2$ to probe the habitability of exoplanets. Their detection will be an enthralling challenge for the future JWST and ARIEL telescopes.

astro-ph.EP

Impact of non-normal error distributions on the benchmarking and ranking of Quantum Machine Learning models

Quantum machine learning models have been gaining significant traction within atomistic simulation communities. Conventionally, relative model performances are being assessed and compared using learning curves (prediction error vs. training set size). This article illustrates the limitations of using the Mean Absolute Error (MAE) for benchmarking, which is particularly relevant in the case of non-normal error distributions. We analyze more specifically the prediction error distribution of the kernel ridge regression with SLATM representation and L 2 distance metric (KRR-SLATM-L2) for effective atomization energies of QM7b molecules calculated at the level of theory CCSD(T)/cc-pVDZ. Error distributions of HF and MP2 at the same basis set referenced to CCSD(T) values were also assessed and compared to the KRR model. We show that the true performance of the KRR-SLATM-L2 method over the QM7b dataset is poorly assessed by the Mean Absolute Error, and can be notably improved after adaptation of the learning set.

physics.data-an

Acknowledging user requirements for accuracy in computational chemistry benchmarks

Computational chemistry has become an important complement to experimental measurements. In order to choose among the multitude of the existing approximations, it is common to use benchmark data sets, and to issue recommendations based on numbers such as mean absolute errors. We argue, using as an example band gaps calculated with density functional approximations, that a more careful study of the benchmark data is needed, stressing that the user's requirements play a role in the choice of an appropriate method. We also appeal to those who measure data capable of being used as a reference, to publish error estimates. We show how the latter can affect the judgment of approximations used in computational chemistry.

physics.chem-ph

Probabilistic performance estimators for computational chemistry methods: Systematic Improvement Probability and Ranking Probability Matrix. II. Applications

In the first part of this study (Paper I), we introduced the systematic improvement probability (SIP) as a tool to assess the level of improvement on absolute errors to be expected when switching between two computational chemistry methods. We developed also two indicators based on robust statistics to address the uncertainty of ranking in computational chemistry benchmarks: Pinv , the inversion probability between two values of a statistic, and Pr , the ranking probability matrix. In this second part, these indicators are applied to nine data sets extracted from the recent benchmarking literature. We illustrate also how the correlation between the error sets might contain useful information on the benchmark dataset quality, notably when experimental data are used as reference.

physics.chem-ph