Search arXivSearch

arXiv · 2605.29976

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

Abstract

We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather forecasting and evaluated up to a 10-day lead time. ArchesWeather is a deterministic model, while ArchesWeatherGen is a probabilistic flow-matching model leveraging ArchesWeather's forecasts, enabling ensemble-based uncertainty quantification. In this work, we adapt these models to act as forced atmospheric models by using additional conditioning on the monthly mean sea surface temperature (SST) and sea ice cover (SIC) as boundary conditions. In particular, we follow the AI Model Intercomparison Project (AIMIP) Phase 1 protocol, which, analogous to the Atmospheric Model Intercomparison Project (AMIP), proposes a standardized experimental setup to evaluate the climate skill of ML-based forced atmospheric models. We present a comprehensive evaluation of both models under these conditions, including comparison against numerical climate models, ablation studies that examine key design choices in the extension, and an analysis of forced versus unforced configurations. Despite being originally developed for weather forecasting, we demonstrate that forced configurations of ArchesWeather and ArchesWeatherGen produce stable long-term climate simulations, have a stable annual cycle, and capture the drift of many climate variables. The models faithfully reproduce ERA5's climatology, large-scale circulations and interannual variability, and they capture the tails of the distributions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Renu Singh, Robert Brunstein, Antonia Jost, Yana Hasson, Thomas Rackow, Claire Monteleoni, Christian Lessig, Guillaume Couairon. 2026-08-18. Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations. https://arxiv.org/abs/2605.29976

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Surface Stokes drift from compact drifting wave buoys

Surface Stokes drift depends strongly on the energy and directions of short waves, which are incompletely resolved by routine wave observations. We derive surface Stokes drift vectors from wave measurements collected by compact drifting buoys during three deployments in the North-East Atlantic and the Alboran Sea. The calculation uses vertical-acceleration spectra and first directional Fourier moments, which describe the mean wave direction and directional concentration at each frequency; it accounts for the Doppler shift caused by buoy motion relative to the water and adds a calibrated high-frequency tail above an intrinsic frequency of 0.7 Hz. Across 13,139 records, the median estimated speed is 0.081 m/s at a median wind speed of 6.8 m/s. Over the measured band of 0.04-1 Hz, accounting for wave directions reduces the magnitude by a median 39% relative to the unidirectional assumption. The median ratio of the parameterised tail magnitude above 0.7 Hz to the total estimated magnitude is 0.37. Comparisons with WAVEWATCH III and Copernicus Marine MFWAM show strong covariation and similar wind-dependent differences from the buoy-derived estimates. On the station-matched sample from the two Atlantic deployments, WAVEWATCH III directional spectra indicate that these differences within the compared band arise mainly from spectral levels rather than from net directional reduction. The observations provide constraints for model evaluation; the contribution of the unresolved short waves remains sensitive to the assumed spectral tail and its directional spreading.

physics.ao-ph

Diffusion-Based Super-Resolution of Adriatic Sea Oceanographic Fields

High-resolution oceanographic fields are critical for resolving mesoscale and sub-mesoscale coastal dynamics, yet their generation remains constrained by both computational cost and observational sparsity. We present OcDiffSR, a conditional denoising diffusion probabilistic model (DDPM) for oceanographic super-resolution that reconstructs high-resolution sea-surface fields from coarse-resolution reanalysis inputs. The model is trained on ten years (2011-2020) of paired low-resolution (GLORYS12V1, 1/12) and high-resolution (Mediterranean Sea Physics Reanalysis, Med MFC, 1/24) data, and evaluated on an independent test year (2009) over the Adriatic Sea. OcDiffSR employs a conditional U-Net augmented with multi-scale low-resolution encoders, cross-attention bottleneck layers, and sinusoidal seasonal embeddings via Feature-wise Linear Modulation (FiLM), enabling joint super-resolution of sea-surface temperature (SST), salinity (SSS), and horizontal velocity components with visually coherent circulation patterns. Benchmarked against bilinear interpolation and the state-of-the-art residual diffusion model CorrDiff, OcDiffSR achieves substantially lower reconstruction errors for scalar fields (RMSESST=0.477 C, RMSESSS=0.346 psu), near-unity Pearson correlation (PCC >= 0.999), and high structural similarity (SSIM >= 0.964). For dynamical vector fields, OcDiffSR outperforms both baselines in absolute error and spatial coherence, though moderate correlation (PCC = 0.64) reflects the intrinsic stochasticity of oceanic velocity fields. Daily and monthly evaluations confirm temporal robustness across all seasons. These results establish OcDiffSR as a reliable framework for high-fidelity oceanographic downscaling and reanalysis enhancement, producing fields that are visually consistent with known ocean dynamics.

physics.ao-ph

Unreported large errors from two PAMGuard three-dimensional localizers of whale calls

Confidence intervals of location (CIL) of calling marine mammals, derived from time-differences-of-arrival (TDOA) between receivers, depend on errors of TDOAs, receiver location, clocks, sound speeds, and location method. When these errors are minuscule, simulations yield small errors of PAMGuard's 3D simplex localizer when click sounds of beaked and sperm whales originate in a 1000 x 1000 x 1000 $\mbox{m}^3$ region using five receivers having horizontal and vertical separations of 1000 m and 150 m respectively. Realistic uncertainties of sound speed up to $\pm 10$ m/s lead to errors up to $10^{14}$ m. With clocks maintained by atomic standards and common practice of correcting TDOA from synchronization measurements at the start and end of an experiment, errors of location are up to $10^{4}$ m. Errors up to $10^2$ and $10^3$ m are found when the receiver's locations are uncertain within 10 and 40 m respectively. Errors of PAMGuard's 3D hyperbolic localizer are almost independent of the above uncertainties, yielding errors of location up to about $10^4$ m even when simulated errors are minuscule. Causes of PAMGuard's 3D location errors are unknown. These algorithms are briefly compared to another method designed to yield a reliable CIL.

physics.ao-ph