Search arXivSearch

arXiv · 2510.25045

Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model

Abstract

Recent advances in AI-based weather prediction have led to the development of artificial intelligence weather prediction (AIWP) models with competitive forecast skill compared to traditional NWP models, but with substantially reduced computational cost. There is a strong need for appropriate methods to evaluate their ability to predict extreme weather events, particularly when spatial coherence is important, and grid resolutions differ between models. We introduce a verification framework that combines spatial verification methods and weighted proper scoring rules. Specifically, the framework extends the High-Resolution Assessment (HiRA) approach with threshold-weighted scoring rules. It enables user-oriented evaluation consistent with how forecasts may be interpreted by operational meteorologists or used in simple post-processing systems. The method supports targeted evaluation of extreme events by allowing flexible weighting of the relative importance of different decision thresholds. We demonstrate this framework by evaluating 32 months of precipitation forecasts from an AIWP model and a high-resolution NWP model. Our results show that model rankings are sensitive to the choice of neighbourhood size. Increasing the neighbourhood size has a greater impact on scores evaluating extreme-event performance for the high-resolution NWP model than for the AIWP model. At near equivalent neighbourhood sizes, the empirical CDF of the high-resolution NWP model only outperformed the empirical CDF of the AIWP model in predicting extreme precipitation events at short lead times. We also demonstrate how this approach can be extended to evaluate discrimination ability in predicting heavy precipitation. We find that the high-resolution NWP model had superior discrimination ability at short lead times.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nicholas Loveday, Tracy Hertneky. 2026-09-08. Evaluating Extreme Precipitation Forecasts: A Threshold-Weighted, Spatial Verification Approach for Comparing an AI Weather Prediction Model Against a High-Resolution NWP Model. https://arxiv.org/abs/2510.25045

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Surface Stokes drift from compact drifting wave buoys

Surface Stokes drift depends strongly on the energy and directions of short waves, which are incompletely resolved by routine wave observations. We derive surface Stokes drift vectors from wave measurements collected by compact drifting buoys during three deployments in the North-East Atlantic and the Alboran Sea. The calculation uses vertical-acceleration spectra and first directional Fourier moments, which describe the mean wave direction and directional concentration at each frequency; it accounts for the Doppler shift caused by buoy motion relative to the water and adds a calibrated high-frequency tail above an intrinsic frequency of 0.7 Hz. Across 13,139 records, the median estimated speed is 0.081 m/s at a median wind speed of 6.8 m/s. Over the measured band of 0.04-1 Hz, accounting for wave directions reduces the magnitude by a median 39% relative to the unidirectional assumption. The median ratio of the parameterised tail magnitude above 0.7 Hz to the total estimated magnitude is 0.37. Comparisons with WAVEWATCH III and Copernicus Marine MFWAM show strong covariation and similar wind-dependent differences from the buoy-derived estimates. On the station-matched sample from the two Atlantic deployments, WAVEWATCH III directional spectra indicate that these differences within the compared band arise mainly from spectral levels rather than from net directional reduction. The observations provide constraints for model evaluation; the contribution of the unresolved short waves remains sensitive to the assumed spectral tail and its directional spreading.

physics.ao-ph

Unreported large errors from two PAMGuard three-dimensional localizers of whale calls

Confidence intervals of location (CIL) of calling marine mammals, derived from time-differences-of-arrival (TDOA) between receivers, depend on errors of TDOAs, receiver location, clocks, sound speeds, and location method. When these errors are minuscule, simulations yield small errors of PAMGuard's 3D simplex localizer when click sounds of beaked and sperm whales originate in a 1000 x 1000 x 1000 $\mbox{m}^3$ region using five receivers having horizontal and vertical separations of 1000 m and 150 m respectively. Realistic uncertainties of sound speed up to $\pm 10$ m/s lead to errors up to $10^{14}$ m. With clocks maintained by atomic standards and common practice of correcting TDOA from synchronization measurements at the start and end of an experiment, errors of location are up to $10^{4}$ m. Errors up to $10^2$ and $10^3$ m are found when the receiver's locations are uncertain within 10 and 40 m respectively. Errors of PAMGuard's 3D hyperbolic localizer are almost independent of the above uncertainties, yielding errors of location up to about $10^4$ m even when simulated errors are minuscule. Causes of PAMGuard's 3D location errors are unknown. These algorithms are briefly compared to another method designed to yield a reliable CIL.

physics.ao-ph

Tropospheric Ozone Formation Potential and Related Design Considerations for Radiative Coolers

Recently, radiative coolers have been widely explored for reducing cooling loads or lowering temperatures in buildings, and at urban scales as a heat-mitigation measure. However, the potential impacts of radiative cooler deployment on the chemical composition of the atmosphere remain largely unexplored. A defining feature of recently-designed radiative coolers is their high ultraviolet (UV) reflectance, which is required for sub-ambient cooling under strong sunlight. Yet, wide adoption of such UV-reflective radiative coolers could substantially increase the UV actinic flux in the atmosphere above. This, in turn, may affect tropospheric ozone concentrations, particularly in urban atmospheres with high NOx concentrations. Here, as a case study, we use a 0-dimensional photochemical box model, constrained by field measurements of meteorological conditions and chemical concentrations in the urban environment of Houston, Texas, to explore the potential impact of the widespread use of UV-reflective radiative coolers on ozone concentrations. Our calculations show that complete deployment of radiative coolers may increase tropospheric ozone levels by as much as 30% during specific meteorological conditions in Houston. Informed by the wavelength-dependent modelling results, we propose specific designs, namely pigmented radiative coolers with different UV reflectances, and UV-absorptive visible-reemitting fluorescent radiative coolers, that could minimize negative ozone formation while retaining appreciable cooling performance. Our results motivate further study on the effects of widespread deployment of radiative cooling designs like superwhite roof paints on air quality, and materials that simultaneously minimize adverse photochemical impact and maximize cooling performance.

physics.ao-ph