Search arXivSearch

arXiv subjects

Eric Cator

Publications and source records attributed to Eric Cator.

At least 19 recordsLinked to original sources

Gibbs conditioning principle for log-concave independent random variables

Let $\nu_1,\nu_2,\dots$ be a sequence of probabilities on the nonnegative integers, and $X=(X_1,X_2, \dots)$ be a sequence of independent random variables $X_i$ with law $\nu_i$. For $\lambda>0$ denote $Z^\lambda_i:= \sum_x \lambda^x\nu_i(x)$ and $\lambda^{\max}:= \sup\{\lambda>0: Z^\lambda_i<\infty \text{ for all }i\}$, and assume $\lambda^{\max}>1$. For $\lambda<\lambda^{\max}$, define the tilted probability $\nu_i^{\lambda}(x):= \lambda^x\nu_i(x)/Z^{\lambda}_i$, and let $X^\lambda$ be a sequence of independent variables $X^\lambda_i$ with law $\nu^{\lambda}_i$, and denote $S^\lambda_n:= X^{\lambda}_1+\dots+X^{\lambda}_n$, with $S_n=S^1_n$. Choose $\lambda^*\in(1,\lambda^{\max})$ and denote $R^*_n:= E (S^{\lambda^*}_n)$. The Gibbs Conditioning Principle (GCP) holds if $P(X\in\cdot|S_n>R^*_n)$ converges weakly to the law of $X^{\lambda^*}$, as $n\to\infty$. We prove the GCP for log-concave $\nu_i$'s, meaning $\nu_i(x+1)\,\nu_i(x-1) \le ( \nu_i(x))^2$, subject to a technical condition that prevents condensation. The canonical measures are the distributions of the first $n$ variables, conditioned on their sum being $k$. Efron's theorem states that for log-concave $\nu_i$'s, the canonical measures are stochastically ordered with respect to $k$. This, in turn, leads to the ordering of the conditioned tilted measures $P(X^\lambda\in\cdot|S^\lambda_n>R^*_n)$ in terms of $\lambda$. This ordering is a fundamental component of our proof.

math.PR

Two-Stage Testing in a high dimensional setting

In a high dimensional regression setting in which the number of variables ($p$) is much larger than the sample size ($n$), the number of possible two-way interactions between the variables is immense. If the number of variables is in the order of one million, which is usually the case in e.g., genetics, the number of two-way interactions is of the order one million squared. In the pursuit of detecting two-way interactions, testing all pairs for interactions one-by-one is computational unfeasible and the multiple testing correction will be severe. In this paper we describe a two-stage testing procedure consisting of a screening and an evaluation stage. It is proven that, under some assumptions, the tests-statistics in the two stages are asymptotically independent. As a result, multiplicity correction in the second stage is only needed for the number of statistical tests that are actually performed in that stage. This increases the power of the testing procedure. Also, since the testing procedure in the first stage is computational simple, the computational burden is lowered. Simulations have been performed for multiple settings and regression models (generalized linear models and Cox PH model) to study the performance of the two-stage testing procedure. The results show type I error control and an increase in power compared to the procedure in which the pairs are tested one-by-one.

stat.ME

Composite Quantile Regression With XGBoost Using the Novel Arctan Pinball Loss

This paper explores the use of XGBoost for composite quantile regression. XGBoost is a highly popular model renowned for its flexibility, efficiency, and capability to deal with missing data. The optimization uses a second order approximation of the loss function, complicating the use of loss functions with a zero or vanishing second derivative. Quantile regression -- a popular approach to obtain conditional quantiles when point estimates alone are insufficient -- unfortunately uses such a loss function, the pinball loss. Existing workarounds are typically inefficient and can result in severe quantile crossings. In this paper, we present a smooth approximation of the pinball loss, the arctan pinball loss, that is tailored to the needs of XGBoost. Specifically, contrary to other smooth approximations, the arctan pinball loss has a relatively large second derivative, which makes it more suitable to use in the second order approximation. Using this loss function enables the simultaneous prediction of multiple quantiles, which is more efficient and results in far fewer quantile crossings.

stat.ML

Likelihood-ratio-based confidence intervals for neural networks

This paper introduces a first implementation of a novel likelihood-ratio-based approach for constructing confidence intervals for neural networks. Our method, called DeepLR, offers several qualitative advantages: most notably, the ability to construct asymmetric intervals that expand in regions with a limited amount of data, and the inherent incorporation of factors such as the amount of training time, network architecture, and regularization techniques. While acknowledging that the current implementation of the method is prohibitively expensive for many deep-learning applications, the high cost may already be justified in specific fields like medical predictions or astrophysics, where a reliable uncertainty estimate for a single prediction is essential. This work highlights the significant potential of a likelihood-ratio-based uncertainty estimate and establishes a promising avenue for future research.

stat.ML

Consistency Tests for Comparing Astrophysical Models and Observations

In astronomy, there is an opportunity to enhance the practice of validating models through statistical techniques, specifically to account for measurement error uncertainties. While models are commonly used to describe observations, there are instances where there is a lack of agreement between the two. This can occur when models are derived from incomplete theories, when a better-fitting model is not available or when measurement uncertainties are not correctly considered. However, with the application of specific tests that assess the consistency between observations and astrophysical models in a model-independent way, it is possible to address this issue. The consistency tests (ConTESTs) developed in this paper use a combination of non-parametric methods and distance measures to obtain a test statistic that evaluates the closeness of the astrophysical model to the observations. To draw conclusions on the consistency hypothesis, a simulation-based methodology is performed. In particular, we built two tests for density models and two for regression models to be used depending on the case at hand and the power of the test needed. We used ConTEST to examine synthetic examples in order to determine the effectiveness of the tests and provide guidance on using them while building a model. We also applied ConTEST to various astronomy cases, identifying which models were consistent and, if not, identifying the probable causes of rejection.

astro-ph.IM

Optimal Training of Mean Variance Estimation Neural Networks

This paper focusses on the optimal implementation of a Mean Variance Estimation network (MVE network) (Nix and Weigend, 1994). This type of network is often used as a building block for uncertainty estimation methods in a regression setting, for instance Concrete dropout (Gal et al., 2017) and Deep Ensembles (Lakshminarayanan et al., 2017). Specifically, an MVE network assumes that the data is produced from a normal distribution with a mean function and variance function. The MVE network outputs a mean and variance estimate and optimizes the network parameters by minimizing the negative loglikelihood. In our paper, we present two significant insights. Firstly, the convergence difficulties reported in recent work can be relatively easily prevented by following the simple yet often overlooked recommendation from the original authors that a warm-up period should be used. During this period, only the mean is optimized with a fixed variance. We demonstrate the effectiveness of this step through experimentation, highlighting that it should be standard practice. As a sidenote, we examine whether, after the warm-up, it is beneficial to fix the mean while optimizing the variance or to optimize both simultaneously. Here, we do not observe a substantial difference. Secondly, we introduce a novel improvement of the MVE network: separate regularization of the mean and the variance estimate. We demonstrate, both on toy examples and on a number of benchmark UCI regression data sets, that following the original recommendations and the novel separate regularization can lead to significant improvements.

stat.ML

Scenario Parameter Generation Method and Scenario Representativeness Metric for Scenario-Based Assessment of Automated Vehicles

The development of assessment methods for the performance of Automated Vehicles (AVs) is essential to enable the deployment of automated driving technologies, due to the complex operational domain of AVs. One candidate is scenario-based assessment, in which test cases are derived from real-world road traffic scenarios obtained from driving data. Because of the high variety of the possible scenarios, using only observed scenarios for the assessment is not sufficient. Therefore, methods for generating additional scenarios are necessary. Our contribution is twofold. First, we propose a method to determine the parameters that describe the scenarios to a sufficient degree without relying on strong assumptions on the parameters that characterize the scenarios. By estimating the probability density function (pdf) of these parameters, realistic parameter values can be generated. Second, we present the Scenario Representativeness (SR) metric based on the Wasserstein distance, which quantifies to what extent the scenarios with the generated parameter values are representative of real-world scenarios while covering the actual variety found in the real-world scenarios. A comparison of our proposed method with methods relying on assumptions of the scenario parametrization and pdf estimation shows that the proposed method can automatically determine the optimal scenario parametrization and pdf estimation. Furthermore, we demonstrate that our SR metric can be used to choose the (number of) parameters that best describe a scenario. The presented method is promising, because the parameterization and pdf estimation can directly be applied to already available importance sampling strategies for accelerating the evaluation of AVs.

cs.RO

Confident Neural Network Regression with Bootstrapped Deep Ensembles

With the rise of the popularity and usage of neural networks, trustworthy uncertainty estimation is becoming increasingly essential. One of the most prominent uncertainty estimation methods is Deep Ensembles (Lakshminarayanan et al., 2017) . A classical parametric model has uncertainty in the parameters due to the fact that the data on which the model is build is a random sample. A modern neural network has an additional uncertainty component since the optimization of the network is random. Lakshminarayanan et al. (2017) noted that Deep Ensembles do not incorporate the classical uncertainty induced by the effect of finite data. In this paper, we present a computationally cheap extension of Deep Ensembles for the regression setting, called Bootstrapped Deep Ensembles, that explicitly takes this classical effect of finite data into account using a modified version of the parametric bootstrap. We demonstrate through an experimental study that our method significantly improves upon standard Deep Ensembles

stat.ML

AutoSourceID-Light. Fast Optical Source Localization via U-Net and Laplacian of Gaussian

$\textbf{Aims}$. With the ever-increasing survey speed of optical wide-field telescopes and the importance of discovering transients when they are still young, rapid and reliable source localization is paramount. We present AutoSourceID-Light (ASID-L), an innovative framework that uses computer vision techniques that can naturally deal with large amounts of data and rapidly localize sources in optical images. $\textbf{Methods}$. We show that the AutoSourceID-Light algorithm based on U-shaped networks and enhanced with a Laplacian of Gaussian filter (Chen et al. 1987) enables outstanding performances in the localization of sources. A U-Net (Ronneberger et al. 2015) network discerns the sources in the images from many different artifacts and passes the result to a Laplacian of Gaussian filter that then estimates the exact location. $\textbf{Results}$. Application on optical images of the MeerLICHT telescope demonstrates the great speed and localization power of the method. We compare the results with the widely used SExtractor (Bertin & Arnouts 1996) and show the out-performances of our method. AutoSourceID-Light rapidly detects more sources not only in low and mid crowded fields, but particularly in areas with more than 150 sources per square arcminute.

astro-ph.IM

The GW-Universe Toolbox II: constraining the binary black hole population with second and third generation detectors

We employ the method used by the GW-Universe Toolbox to generate a synthetic catalogue of detection of stellar mass binary black hole (BBH) mergers. We study advanced LIGO (aLIGO) and Einstein Telescope (ET) as two representatives for the 2nd and 3rd generation GW observatories, and study how GW observations of BBHs can be used to constrain the merger rate as function of redshift and masses. We also simulate the observations from a detector that is half as sensitive as the ET at design which represents an early phase of ET. Two methods are used to obtain the constraints on the source population properties from the catalogues: 1. parametric differential merger rate model and applies a Bayesian inference on the parameters; and 2. non-parametric and uses weighted Kernel density estimators. The results show the overwhelming advantages of the 3rd generation detector over the 2nd generation for the study of BBH population properties, especially at a redshifts higher than ~2, where the merger rate is believed to peak. With the simulated aLIGO catalogue, the parametric Bayesian method can still give some constraints on the merger rate density and mass function beyond its detecting horizon, while the non-parametric method lose the constraining ability completely there. We also find that, despite the numbers of detection of the half-ET can be easily compatible with full ET after a longer observation duration, the catalogue from the full ET can still give much better constraints on the population properties, due to its smaller uncertainties on the physical parameters of the GW events.

astro-ph.HE

Constrained Sampling from a Kernel Density Estimator to Generate Scenarios for the Assessment of Automated Vehicles

The safety assessment of automated vehicles (AVs) is an important aspect of the development cycle of AVs. A scenario-based assessment approach is accepted by many players in the field as part of the complete safety assessment. A scenario is a representation of a situation on the road to which the AV needs to respond appropriately. One way to generate the required scenario-based test descriptions is to parameterize the scenarios and to draw these parameters from a probability density function (pdf). Because the shape of the pdf is unknown beforehand, assuming a functional form of the pdf and fitting the parameters to the data may lead to inaccurate fits. As an alternative, Kernel Density Estimation (KDE) is a promising candidate for estimating the underlying pdf, because it is flexible with the underlying distribution of the parameters. Drawing random samples from a pdf estimated with KDE is possible without the need of evaluating the actual pdf, which makes it suitable for drawing random samples for, e.g., Monte Carlo methods. Sampling from a KDE while the samples satisfy a linear equality constraint, however, has not been described in the literature, as far as the authors know. In this paper, we propose a method to sample from a pdf estimated using KDE, such that the samples satisfy a linear equality constraint. We also present an algorithm of our method in pseudo-code. The method can be used to generating scenarios that have, e.g., a predetermined starting speed or to generate different types of scenarios. This paper also shows that the method for sampling scenarios can be used in case a Singular Value Decomposition (SVD) is used to reduce the dimension of the parameter vectors.

cs.AI

How to Evaluate Uncertainty Estimates in Machine Learning for Regression?

As neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by evaluating their loglikelihood on a test set. Other methods output a prediction interval directly. These methods are often tested by examining the fraction of test points that fall inside the corresponding prediction intervals. Intuitively both approaches seem logical. However, we demonstrate through both theoretical arguments and simulations that both ways of evaluating the quality of uncertainty estimates have serious flaws. Firstly, both approaches cannot disentangle the separate components that jointly create the predictive uncertainty, making it difficult to evaluate the quality of the estimates of these components. Secondly, a better loglikelihood does not guarantee better prediction intervals, which is what the methods are often used for in practice. Moreover, the current approach to test prediction intervals directly has additional flaws. We show why it is fundamentally flawed to test a prediction or confidence interval on a single test set. At best, marginal coverage is measured, implicitly averaging out overconfident and underconfident predictions. A much more desirable property is pointwise coverage, requiring the correct coverage for each prediction. We demonstrate through practical examples that these effects can result in favoring a method, based on the predictive uncertainty, that has undesirable behaviour of the confidence or prediction intervals. Finally, we propose a simulation-based testing approach that addresses these problems while still allowing easy comparison between different methods.

stat.ML

Extending the Mann-Kendall test to allow for measurement uncertainty

The Mann-Kendall test for trend has gained a lot of attention in a range of disciplines, especially in the environmental sciences. One of the drawbacks of the Mann-Kendall test when applied to real data is that no distinction can be made between meaningful and non-meaningful differences in subsequent observations. We introduce the concept of partial ties, which allows inferences while accounting for (non)meaningful difference. We introduce the modified statistic that accounts for such a concept and derive its variance estimator. We also present analytical results for the behavior of the test in a class of contiguous alternatives. Simulation results which illustrate the added value of the test are presented. We apply our extended version of the test to some real data concerning blood donation in Europe.

stat.ME

The Significance Filter, the Winner's Curse and the Need to Shrink

The "significance filter" refers to focusing exclusively on statistically significant results. Since frequentist properties such as unbiasedness and coverage are valid only before the data have been observed, there are no guarantees if we condition on significance. In fact, the significance filter leads to overestimation of the magnitude of the parameter, which has been called the "winner's curse". It can also lead to undercoverage of the confidence interval. Moreover, these problems become more severe if the power is low. While these issues clearly deserve our attention, they have been studied only informally and mathematical results are lacking. Here we study them from the frequentist and the Bayesian perspective. We prove that the relative bias of the magnitude is a decreasing function of the power and that the usual confidence interval undercovers when the power is less than 50%. We conclude that failure to apply the appropriate amount of shrinkage can lead to misleading inferences.

stat.ME

On the Geometry of the Last Passage Percolation Problem

We analyze the geometrical structure of the passage times in the last passage percolation model. Viewing the passage time as a piecewise linear function of the weights we determine the domains of the various pieces, which are the subsets of the weight space that make a given path the longest one. We focus on the case when all weights are assumed to be positive, and as a result each domain is a pointed polyhedral cone. We determine the extreme rays, facets, and two-dimensional faces of each cone, and also review a well-known simplicial decomposition of the maximal cones via the so-called order cone. All geometric properties are derived using arguments phrased in terms of the last passage model itself. Our motivation is to understand path probabilities of the extremal corner paths on boxes in $\Z^2$, but all of our arguments apply to general, finite partially ordered sets.

math.PR

Explicit bounds for critical infection rates and expected extinction times of the contact process on finite random graphs

We introduce a method to prove metastability of the contact process on Erd\H{o}s-R\'enyi graphs and on configuration model graphs. The method relies on uniformly bounding the total infection rate from below, over all sets with a fixed number of nodes. Once this bound is established, a simple comparison with a well chosen birth-and-death process will show the exponential growth of the extinction time. Our paper complements recent results on the metastability of the contact process: under a certain minimal edge density condition, we give explicit lower bounds on the infection rate needed to get metastability, and we have explicit exponentially growing lower bounds on the expected extinction time.

math.PR

A new look at distances and velocities of neutron stars

We take a fresh look at the determination of distances and velocities of neutron stars. The conversion of a parallax measurement into a distance, or distance probability distribution, has led to a debate quite similar to the one involving Cepheids, centering on the question whether priors can be used when discussing a single system. With the example of PSRJ0218+4232 we show that a prior is necessary to determine the probability distribution for the distance. The distance of this pulsar implies a gamma-ray luminosity larger than 10% of its spindown luminosity. For velocities the debate is whether a single Maxwellian describes the distribution for young pulsars. By limiting our discussion to accurate (VLBI) measurements we argue that a description with two Maxwellians, with distribution parameters sigma1=77 and sigma2=320 km/s, is significantly better. Corrections for galactic rotation, to derive velocities with respect to the local standards of rest, are insignificant.

astro-ph.HE

The observed velocity distribution of young pulsars

We argue that comparison with observations of theoretical models for the velocity distribution of pulsars must be done directly with the observed quantities, i.e. parallax and the two components of proper motion. We develop a formalism to do so, and apply it to pulsars with accurate VLBI measurements. We find that a distribution with two maxwellians improves significantly on a single maxwellian. The `mixed' model takes into account that pulsars move away from their place of birth, a narrow region around the galactic plane. The best model has 42% of the pulsars in a maxwellian with average velocity sigma sqrt{8/pi}=120 km/s, and 58% in a maxwellian with average velocity 540 km/s. About 5% of the pulsars has a velocity at birth less than 60\,km/s. For the youngest pulsars (tau_c<10 Myr), these numbers are 32% with 130 km/s, 68% with 520 km/s, and 3%, with appreciable uncertainties.

astro-ph.HE