Search arXivSearch

arXiv subjects

Marco Oesting

Publications and source records attributed to Marco Oesting.

At least 19 recordsLinked to original sources

Central limit theory for serial tail dependence estimators in heavy-tailed long memory linear time series

We prove multiple central limit theorems for serial tail dependence estimators in heavy-tailed long memory linear time series. The main theoretical tools are two novel multivariate reduction principles for partial sums of heavy-tailed long memory linear time series, subordinated over sliding windows and above a threshold growing with sample size. This requires addressing several substantial difficulties, including handling a nonlinear, sample-size dependent, and multivariate subordination mechanism, the dependence between several overlapping linear processes, and the lack of higher-order moments of the marginal distribution. Despite these obstacles, our assumptions are mild and, in particular, the innovation process is allowed to have infinite variance. A key feature of our theory is that our second reduction principle holds uniformly in the threshold, allowing central limit theory for empirical extremograms with sample quantiles as thresholds. This question has received little attention in the literature on serial extremal dependence estimation even though the version of empirical extremograms with random thresholds is ubiquitous in practice. We compare our results in several respects with those that may be obtained under short-range dependence, thereby discovering markedly different convergence rates and limit laws in our long memory setting.

math.ST

Asymptotic behavior of spatio-temporal point processes of exceedances

In this paper, we analyze the asymptotic behavior of the point process of exceedances in a spatio-temporal setting whose points are given by the rescaled occurrence times, the sites and the rescaled values of exceedances. Here, the exceedances over a high threshold are flexibly defined via site-dependent risk functionals. Exploiting the framework of stationary regularly varying multivariate time series, we merge and extend the results from the literature in order to show weak convergence of the considered point processes of extremes and to explicitly determine its limit distribution.

math.PR

Bayesian inference for functional extreme events defined via partially unobserved processes

In order to describe the extremal behaviour of some stochastic process $X$, approaches from univariate extreme value theory are typically generalized to the spatial domain. In particular, generalized peaks-over-threshold approaches allow for the consideration of single extreme events. These can be flexibly defined as exceedances of a risk functional $r$, such as a spatial average, applied to $X$. Inference for the resulting limit process, the so-called $r$-Pareto process, requires the evaluation of $r(X)$ and thus the knowledge of the whole process $X$. In many practical applications, however, observations of $X$ are only available at scattered sites. To overcome this issue, we propose a two-step MCMC-algorithm in a Bayesian framework. In a first step, we sample from $X$ conditionally on the observations in order to evaluate which observations lead to $r$-exceedances. In a second step, we use these exceedances to sample from the posterior distribution of the parameters of the limiting $r$-Pareto process. Alternating these steps results in a full Bayesian model for the extremes of $X$. We show that, under appropriate assumptions, the probability of classifying an observation as $r$-exceedance in the first step converges to the desired probability. Furthermore, given the first step, the distribution of the Markov chain constructed in the second step converges to the posterior distribution of interest. The procedure is compared to the Bayesian version of the standard procedure in a simulation study.

stat.ME

Accuracy estimation of neural networks by extreme value theory

Neural networks are able to approximate any continuous function on a compact set. However, it is not obvious how to quantify the error of the neural network, i.e., the remaining bias between the function and the neural network. Here, we propose the application of extreme value theory to quantify large values of the error, which are typically relevant in applications. The distribution of the error beyond some threshold is approximately generalized Pareto distributed. We provide a new estimator of the shape parameter of the Pareto distribution suitable to describe the error of neural networks. Numerical experiments are provided.

stat.ML

Central limit theory for Peaks-over-Threshold partial sums of long memory linear time series

Over the last 30 years, extensive work has been devoted to developing central limit theory for partial sums of subordinated long memory linear time series. A much less studied problem, motivated by questions that are ubiquitous in extreme value theory, is the asymptotic behavior of such partial sums when the subordination mechanism has a threshold depending on sample size, so as to focus on the right tail of the time series. This article substantially extends longstanding asymptotic techniques by allowing the subordination mechanism to depend on the sample size in this way and to grow at a polynomial rate, while permitting the innovation process to have infinite variance. The cornerstone of our theoretical approach is a tailored reduction principle, which enables the use of classical results on partial sums of long memory linear processes. In this way we obtain asymptotic theory for certain Peaks-over-Threshold estimators with deterministic or random thresholds. Applications cover both heavy- and light-tailed regimes, yielding unexpected results which, to the best of our knowledge, are new to the literature. A simulation study illustrates the relevance of our findings in finite samples.

math.PR

Clustering Tails in High Dimension

One potential solution to combat the scarcity of tail observations in extreme value analysis is to integrate information from multiple datasets sharing similar tail properties, for instance, a common extreme value index. In other words, for a multivariate dataset, we intend to group dimensions into clusters first, before applying any pooling techniques. This paper addresses the clustering problem for a high dimensional dataset, according to their extreme value indices. We propose an iterative clustering procedure that sequentially partitions the variables into groups, ordered from the heaviest-tailed to the lightesttailed distributions. At each step, our method identifies and extracts a group of variables that share the highest extreme value index among the remaining ones. This approach differs fundamentally from conventional clustering methods such as using pre-estimated extreme value indices in a two-step clustering method. We show the consistency property of the proposed algorithm and demonstrate its finite-sample performance using a simulation study and a real data application.

stat.ME

Long Memory of Max-Stable Time Series as Phase Transition: Asymptotic Behaviour of Tail Dependence Estimators

In this paper, we consider a simple estimator for tail dependence coefficients of a max-stable time series and show its asymptotic normality under a mild condition. The novelty of our result is that this condition does not involve mixing properties that are common in the literature. More importantly, our condition is linked to the transition between long and short range dependence (LRD/SRD) for max-stable time series. This is based on a recently proposed notion of LRD in the sense of indicators of excursion sets which is meaningfully defined for infinite-variance time series. In particular, we show that asymptotic normality with standard rate of convergence and a function of the sum of tail coefficients as asymptotic variance holds if and only if the max-stable time series is SRD.

math.ST

Extremes in High Dimensions: Methods and Scalable Algorithms

Extreme value theory for univariate and low-dimensional observations has been explored in considerable detail, but the field is still in an early stage regarding high-dimensional settings. This paper focuses on H\"usler-Reiss models, a popular class of models for multivariate extremes similar to multivariate Gaussian distributions, and their domain of attraction. We develop estimators for the model parameters based on score matching, and we equip these estimators with theories and exceptionally scalable algorithms. Simulations and applications to weather extremes demonstrate the fact that the estimators can estimate a large number of parameters reliably and fast; for example, we show that H\"usler-Reiss models with thousands of parameters can be fitted within a couple of minutes on a standard laptop. More generally speaking, our work relates extreme value theory to modern concepts of high-dimensional statistics and convex optimization.

stat.ME

Non-stationary max-stable models with an application to heavy rainfall data

In recent years, parametric models for max-stable processes have become a popular choice for modeling spatial extremes because they arise as the asymptotic limit of rescaled maxima of independent and identically distributed random processes. Apart from few exceptions for the class of extremal-t processes, existing literature mainly focuses on models with stationary dependence structures. In this paper, we propose a novel non-stationary approach that can be used for both Brown-Resnick and extremal-t processes - two of the most popular classes of max-stable processes - by including covariates in the corresponding variogram and correlation functions, respectively. We apply our new approach to extreme precipitation data in two regions in Southern and Northern Germany and compare the results to existing stationary models in terms of Takeuchi's information criterion (TIC). Our results indicate that, for this case study, non-stationary models are more appropriate than stationary ones for the region in Southern Germany. In addition, we investigate theoretical properties of max-stable processes conditional on random covariates. We show that these can result in both asymptotically dependent and asymptotically independent processes. Thus, conditional models are more flexible than classical max-stable models.

stat.ME

Patterns in Spatio-Temporal Extremes

In environmental science applications, extreme events frequently exhibit a complex spatio-temporal structure, which is difficult to describe flexibly and estimate in a computationally efficient way using state-of-art parametric extreme-value models. In this paper, we propose a computationally-cheap non-parametric approach to investigate the probability distribution of temporal clusters of spatial extremes, and study within-cluster patterns with respect to various characteristics. These include risk functionals describing the overall event magnitude, spatial risk measures such as the size of the affected area, and measures representing the location of the extreme event. Under the framework of functional regular variation, we verify the existence of the corresponding limit distributions as the considered events become increasingly extreme. Furthermore, we develop non-parametric estimators for the limiting expressions of interest and show their asymptotic normality under appropriate mixing conditions. Uncertainty is assessed using a multiplier block bootstrap. The finite-sample behavior of our estimators and the bootstrap scheme is demonstrated in a spatio-temporal simulated example. Our methodology is then applied to study the spatio-temporal dependence structure of high-dimensional sea surface temperature data for the southern Red Sea. Our analysis reveals new insights into the temporal persistence, and the complex hydrodynamic patterns of extreme sea temperature events in this region.

stat.ME

Implications of Modeling Seasonal Differences in the Extremal Dependence of Rainfall Maxima

For modeling extreme rainfall, the widely used Brown-Resnick max-stable model extends the concept of the variogram to suit bloc maxima, allowing the explicit modeling of the extremal dependence shown by the spatial data. This extremal dependence stems from the geometrical characteristics of the observed rainfall, which is associated with different meteorological processes and is usually considered to be constant when designing the model for a study. However, depending on the region, this dependence can change throughout the year, as the prevailing meteorological conditions that drive the rainfall generation process change with the season. Therefore, this study analyzes the impact of the seasonal change in extremal dependence for the modeling of annual block maxima in the Berlin-Brandenburg region. For this study, two seasons were considered as proxies for different dominant meteorological conditions: summer for convective rainfall and winter for frontal/stratiform rainfall. Using maxima from both seasons, we compared the skill of a linear model with spatial covariates (that assumed spatial independence) with the skill of a Brown-Resnick max-stable model. This comparison showed a considerable difference between seasons, with the isotropic Brown-Resnick model showing considerable loss of skill for the winter maxima. We conclude that the assumptions commonly made when using the Brown-Resnick model are appropriate for modeling summer (i.e., convective) events, but further work should be done for modeling other types of precipitation regimes.

physics.ao-ph

$L_p$-norm spherical copulas

In this paper we study $L_p$-norm spherical copulas for arbitrary $p \in [1,\infty]$ and arbitrary dimensions. The study is motivated by a conjecture that these distributions lead to a sharp bound for the value of a certain generalized mean difference. We fully characterize conditions for existence and uniqueness of $L_p$-norm spherical copulas. Explicit formulas for their densities and correlation coefficients are derived and the distribution of the radial part is determined. Moreover, statistical inference and efficient simulation are considered.

math.ST

Detection of Long Range Dependence in the Time Domain for (In)Finite-Variance Time Series

Empirical detection of long range dependence (LRD) of a time series often consists of deciding whether an estimate of the memory parameter $d$ corresponds to LRD. Surprisingly, the literature offers numerous spectral domain estimators for $d$ but there are only a few estimators in the time domain. Moreover, the latter estimators are criticized for relying on visual inspection to determine an observation window $[n_1, n_2]$ for a linear regression to run on. Theoretically motivated choices of $n_1$ and $n_2$ are often missing for many time series models. In this paper, we take the well-known variance plot estimator and provide rigorous asymptotic conditions on $[n_1, n_2]$ to ensure the estimator's consistency under LRD. We establish these conditions for a large class of square-integrable time series models. This large class enables one to use the variance plot estimator to detect LRD for infinite-variance time series (after suitable transformation). Thus, detection of LRD for infinite-variance time series is another novelty of our paper. A simulation study indicates that the variance plot estimator can detect LRD better than the popular spectral domain GPH estimator.

math.ST

Evaluation of binary classifiers for asymptotically dependent and independent extremes

Machine learning classification methods usually assume that all possible classes are sufficiently present within the training set. Due to their inherent rarities, extreme events are always under-represented and classifiers tailored for predicting extremes need to be carefully designed to handle this under-representation. In this paper, we address the question of how to assess and compare classifiers with respect to their capacity to capture extreme occurrences. This is also related to the topic of scoring rules used in forecasting literature. In this context, we propose and study a risk function adapted to extremal classifiers. The inferential properties of our empirical risk estimator are derived under the framework of multivariate regular variation and hidden regular variation. A simulation study compares different classifiers and indicates their performance with respect to our risk function. To conclude, we apply our framework to the analysis of extreme river discharges in the Danube river basin. The application compares different predictive algorithms and test their capacity at forecasting river discharges from other river stations.

stat.ME

Estimation of the Spectral Measure from ConvexCombinations of Regularly Varying RandomVectors

The extremal dependence structure of a regularly varying random vector Xis fully described by its limiting spectral measure. In this paper, we investigate how torecover characteristics of the measure, such as extremal coefficients, from the extremalbehaviour of convex combinations of components of X. Our considerations result in aclass of new estimators of moments of the corresponding combinations for the spectralvector. We show asymptotic normality by means of a functional limit theorem and, focusingon the estimation of extremal coefficients, we verify that the minimal asymptoticvariance can be achieved by a plug-in estimator using subsampling bootstrap. We illustratethe benefits of our approach on simulated and real data.

math.ST

Spatial Modeling of Heavy Precipitation by Coupling Weather Station Recordings and Ensemble Forecasts with Max-Stable Processes

Due to complex physical phenomena, the distribution of heavy rainfall events is difficult to model spatially. Physically based numerical models can often provide physically coherent spatial patterns, but may miss some important precipitation features like heavy rainfall intensities. Measurements at ground-based weather stations, however, supply adequate rainfall intensities, but most national weather recording networks are often spatially too sparse to capture rainfall patterns adequately. To bring the best out of these two sources of information, climatologists and hydrologists have been seeking models that can efficiently merge different types of rainfall data. One inherent difficulty is to capture the appropriate multivariate dependence structure among rainfall maxima. For this purpose, multivariate extreme value theory suggests the use of a max-stable process. Such a process can be represented by a max-linear combination of independent copies of a hidden stochastic process weighted by a Poisson point process. In practice, the choice of this hidden process is non-trivial, especially if anisotropy, non-stationarity and nugget effects are present in the spatial data at hand. By coupling forecast ensemble data from the French national weather service (M\'et\'eo-France) with local observations, we construct and compare different types of data driven max-stable processes that are parsimonious in parameters, easy to simulate and capable of reproducing nugget effects and spatial non-stationarities. We also compare our new method with classical approaches from spatial extreme value theory such as Brown-Resnick processes.

stat.AP

Long Range Dependence for Stable Random Processes

We investigate long and short memory in $\alpha$-stable moving averages and max-stable processes with $\alpha$-Fr\'echet marginal distributions. As these processes are heavy-tailed, we rely on the notion of long range dependence suggested by Kulik and Spodarev (2019) based on the covariance of excursions. Sufficient conditions for the long and short range dependence of $\alpha$-stable moving averages are proven in terms of integrability of the corresponding kernel functions. For max-stable processes, the extremal coefficient function is used to state a necessary and sufficient condition for long range dependence.

math.PR

Sampling Sup-Normalized Spectral Functions for Brown-Resnick Processes

Sup-normalized spectral functions form building blocks of max-stable and Pareto processes and therefore play an important role in modeling spatial extremes. For one of the most popular examples, the Brown-Resnick process, simulation is not straightforward. In this paper, we generalize two approaches for simulation via Markov Chain Monte Carlo methods and rejection sampling by introducing new classes of proposal densities. In both cases, we provide an optimal choice of the proposal density with respect to sampling efficiency. The performance of the procedures is demonstrated in an example.

math.ST