Search arXiv⌕ Search

arXiv · 2512.14898

Predicting Forecast Error for the HRRR Using LSTM Neural Networks: A Comparative Study Using New York and Oklahoma State Mesonets

Abstract

Long Short-Term Memory (LSTM) models are trained to predict forecast errors for the High-Resolution Rapid Refresh (HRRR) model using the New York State Mesonet and Oklahoma State Mesonet near-surface weather observations as ground truth. When evaluated using mean-absolute-error and percent improvement relative to the HRRR, LSTMs predict precipitation error most accurately, providing, on average, a 48% improvement relative to the HRRR forecast, followed by wind error, providing, on average, a 15% improvement, and then temperature error, providing, on average, a 25% improvement. Precipitation errors exhibit an asymmetry, with overforecast precipitation detected more accurately than underforecast, while wind error predictions are consistent across over- and underforecast predictions. Temperature error predictions are relatively accurate but smoother, with respect to variance, than true observations. This paper describes an overview of LSTM performance with the expressed intent of providing forecasters with real-time predictions of forecast error at the point of use within the New York State and Oklahoma State Mesonets. In practice, the predicted errors can be used to adjust deterministic HRRR forecasts at the point of use, identify locations and variables with elevated uncertainty, and provide supplemental guidance for high-impact decision-making. This research demonstrates the potential of LSTM-based machine learning models to provide actionable, location-specific predictions of forecast error for high-resolution operational numerical weather prediction (NWP) systems. However, model performance is variable-dependent, and the approach relies on the availability of dense mesonet observations, which may limit applicability in data-sparse regions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

David Aaron Evans, Kara J. Sulia, Nick P. Bassill, Chris D. Thorncroft, Jay C. Rothenberger, Lauriana C. Gaudet. 2026-05-14. Predicting Forecast Error for the HRRR Using LSTM Neural Networks: A Comparative Study Using New York and Oklahoma State Mesonets. https://arxiv.org/abs/2512.14898

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The impact of two-dimensional filtering on SWOT observations with white noise

The Surface Water and Ocean Topography (SWOT) mission provides two-dimensional observations of sea surface height (SSH) at unprecedented spatial resolution, enabling exploration of ocean variability down to scales of $O(10~\mathrm{ km})$. At these scales, however, interpreting SSH variability is challenging because ocean dynamical signals overlap with measurement noise, and their respective spectral signatures are not yet fully understood. Recent analyses of SWOT 2-km posting observations have shown that along-track spectra transition to flatter but still red spectra at scales around 30 km. These flatter portions of the spectra have a power-law-like behavior and spectral slopes of approximately $-1$ or steeper, and their magnitudes and slopes are correlated with SWOT measurement noise magnitude. Here, we investigate the hypothesis that these flatter but still red along-track small-scale spectra can arise from two-dimensional filtering and aliasing of spatially uncorrelated (white) noise. Using synthetic experiments, we show that the resulting one-dimensional along-track spectra exhibit a similar transition to red, power-law-like behavior at scales of 15--50 km, qualitatively similar to spectral behavior reported in SWOT observations. The transition scale and apparent spectral slopes depend on the noise level, its cross-track variability, and the background ocean signal. This finding highlights the importance of carefully accounting for measurement noise and processing effects when interpreting SWOT spectra, and suggests that a white noise model should serve as a baseline null hypothesis for small-scale spectral analyses.

physics.ao-ph↗

Improving global precipitation forecasts with an AI weather model trained on satellite observations

Precipitation forecasts shape decision-making across the global economy, particularly in sectors such as agriculture. However, unlike variables such as temperature, precipitation is highly intermittent and localized, making it difficult to forecast. While recent advances in AI weather prediction systems have enabled them to surpass physical models on globally averaged metrics, improvements in mean error rarely translate to actionable forecasts of severe flooding or dry crop fields. Furthermore, most of these models are trained and evaluated against a reanalysis data product, ERA5, which has well-known biases. Here we retrain AIFS, ECMWF's widely-used, open-source operational 0.25° probabilistic graph-transformer weather model, on satellite-based precipitation observations. Our model, Laxmi, improves global probabilistic accuracy by 19% and resolves systematic distributional biases in ERA5. Specifically, Laxmi reduces drizzle overprediction by 33% for amounts less than 3 mm per day. It also mitigates extreme rainfall underprediction, improving the global 95th percentile Brier skill score by 57%. Across a case study of 10 Indian tropical storms, Laxmi delivered the most accurate forecast of 150 mm event-total precipitation in 7 events, compared to 1 for AIFS and 2 for the leading physical model, IFS. Our results demonstrate that incorporating observation-based precipitation data directly into training can substantially improve forecasts.

physics.ao-ph↗

HClimRep-Ocean: A Global Ocean Emulator on an Unstructured Mesh

Machine-learning (ML) emulators for atmospheric processes have advanced rapidly in recent years, transforming weather forecasting. Although early ML ocean forecasting models now exist, they remain less developed than their atmospheric counterparts. Unlike the atmosphere, much of the ocean's kinetic energy resides in mesoscale eddies whose characteristic spatial scales are approximately an order of magnitude smaller than those of comparable atmospheric features. Moreover, complex coastlines, narrow straits, and ice-covered seas make boundary representation a central challenge that atmospheric models do not face. Consequently, numerical ocean simulations commonly use locally refined or even completely unstructured meshes. However, their data-driven counterparts have so far been built around latitude-longitude grids. We present HClimRep-Ocean, an ocean emulator that operates directly on the native unstructured mesh of FESOM2. The emulator is trained on a 209-year AWI-CM3 control integration and is run without atmospheric forcing, receiving the atmospheric state only at initialisation time, which isolates the predictability carried by the ocean state itself. Skill is strongly field-dependent: for currents, HClimRep-Ocean outperforms every reference at 30 day forecast, whereas for temperature and salinity a damped-anomaly persistence forecast remains the more accurate estimator. This behaviour is physically interpretable: current variability is largely geostrophic and internally generated, whereas sea-surface temperature and salinity fluctuations are driven by atmospheric forcing through weather state. Evaluated independently on the OceanBench benchmark, a reanalysis-trained variant of HClimRep-Ocean achieves the lowest RMSE against GLORYS reanalysis among all assessed systems, confirming the competitiveness of the native-mesh approach.

physics.ao-ph↗