Search arXivSearch

arXiv subjects

Juliette Legrand

Publications and source records attributed to Juliette Legrand.

6 recordsLinked to original sources

Extrapolation of extreme covariates in generalized additive regression using extreme-value theory

Predictions with covariates in the tails of the covariate distribution often lack accuracy and are highly uncertain but may be of critical importance in many applications. New covariates may even lie beyond the range of training data. The problem can be particularly critical in environmental contexts, for example with climate-change scenarios, where new covariates represent future, more extreme climate conditions. We here propose novel methods to improve generalized additive models (GAMs) with focus on extreme covariates and covariate extrapolation. Adopting a random-design setting, we continuously integrate GAMs for the bulk of covariate distributions with models motivated by multivariate extreme value theory for high covariate values. We consider continuous responses but develop also a new approach for binary responses by assuming a continuous latent response variable. Our framework imposes a specific structure for large values of the covariates motivated by extreme value theory: by combining a transformation to a specific marginal scale with an appropriate link function, the continuous response variable depends linearly on the covariates. In an application to occurrences and sizes of large wildfires in Europe for the period 2008-2023, we explore how the new method can improve predictions, especially during extreme conditions, using environmental and meteorological covariates.

stat.ME

Bayesian spatial modelling framework for assessing residential flood risk in property insurance

Spatial heterogeneity in insurance risk modelling is often represented using coarse areal structures, which can obscure fine-scale patterns critical for accurate risk assessment. This study introduces a point-referenced Bayesian framework to model claim occurrence and severity at the policyholder level, avoiding reliance on predefined geographic aggregation. Drawing on a large French insurance portfolio combined with high-resolution environmental variables, rainfall records, and institutional hazard maps, we compare a benchmark GLM with several discrete Bayesian specifications, including independent random effects, intrinsic conditional autoregressive (iCAR) and Besag-York-Mollie (BYM) models, and a continuously indexed Gaussian random field constructed using the stochastic partial differential equation (SPDE) approach. Inference is performed using Integrated Nested Laplace Approximation (INLA), enabling efficient estimation of latent spatial fields and non-linear covariate effects. Our results show that accounting for spatial dependence substantially improves occurrence modelling, while gains in severity prediction are more limited. The SPDE formulation further outperforms areal models by capturing sub-municipal risk gradients and reducing artefacts induced by arbitrary geographic partitioning. By conditioning on detailed building-level attributes, we isolate the contribution of latent spatial effects, refine the interpretation of observed covariates, and improve the allocation of risk premiums across the portfolio. In addition to enhanced predictive performance, the framework provides coherent uncertainty quantification and supports tail-risk assessment. To our knowledge, this is the first application of point-referenced SPDE models to flood insurance, offering a scalable statistical alternative for pricing and managing risks with strong spatial structure.

stat.AP

Contributions of geolocated weather and building related data for insurance assessment of flood risks

Floods rank among the costliest natural hazards, causing over USD 100 billion in insured losses between 2013 and 2023. In France, persistent deficits in the natural catastrophe scheme highlight the need for accurate, building-scale flood risk assessment. Insurers typically rely on frequency-severity models supported by hazard maps and regional climate indicators. However, previous studies show that such large-scale variables explain only a limited share of the variability in individual flood losses. This study evaluates the marginal contribution of multiple georeferenced data layers to modeling flood claim occurrence and severity in a large French home insurance portfolio. Starting from a baseline model based on standard underwriting information, we sequentially introduce climate-expert variables, extreme rainfall indicators, and fine-scale geolocated building and environmental attributes. The analysis focuses on a practical setting in which insurers cannot deploy full hydrological or hydraulic catastrophe models because of budgetary, licensing, or operational constraints. Results show that rainfall-based indicators, particularly a newly constructed metric capturing intense local precipitation, substantially improve claim modeling performance. Building and environmental variables further enhance occurrence prediction. Overall, the findings demonstrate how high-resolution geolocated data improve exposure and vulnerability assessment, complement official flood maps, and provide insurers with an operational framework for refining flood risk evaluation and pricing.

stat.AP

Evaluation of binary classifiers for asymptotically dependent and independent extremes

Machine learning classification methods usually assume that all possible classes are sufficiently present within the training set. Due to their inherent rarities, extreme events are always under-represented and classifiers tailored for predicting extremes need to be carefully designed to handle this under-representation. In this paper, we address the question of how to assess and compare classifiers with respect to their capacity to capture extreme occurrences. This is also related to the topic of scoring rules used in forecasting literature. In this context, we propose and study a risk function adapted to extremal classifiers. The inferential properties of our empirical risk estimator are derived under the framework of multivariate regular variation and hidden regular variation. A simulation study compares different classifiers and indicates their performance with respect to our risk function. To conclude, we apply our framework to the analysis of extreme river discharges in the Danube river basin. The application compares different predictive algorithms and test their capacity at forecasting river discharges from other river stations.

stat.ME

Pareto processes for threshold exceedances in spatial extremes

We review some recent development in the theory of spatial extremes related to Pareto Processes and modeling of threshold exceedances. We provide theoretical background, methodology for modeling, simulation and inference as well as an illustration to wave height modelling. This preprint is an author version of a chapter to appear in a collaborative book.

math.ST

Assessing Extreme Risk using Stochastic Simulation of Extremes

Risk management is particularly concerned with extreme events, but analysing these events is often hindered by the scarcity of data, especially in a multivariate context. This data scarcity complicates risk management efforts. Various tools can assess the risk posed by extreme events, even under extraordinary circumstances. This paper studies the evaluation of univariate risk for a given risk factor using metrics that account for its asymptotic dependence on other risk factors. Data availability is crucial, particularly for extreme events where it is often limited by the nature of the phenomenon itself, making estimation challenging. To address this issue, two non-parametric simulation algorithms based on multivariate extreme theory are developed. These algorithms aim to extend a sample of extremes jointly and conditionally for asymptotically dependent variables using stochastic simulation and multivariate Generalised Pareto Distributions. The approach is illustrated with numerical analyses of both simulated and real data to assess the accuracy of extreme risk metric estimations.

stat.ME