Search arXiv⌕ Search

arXiv · 2610.03172

Domain-Adaptive Data Assimilation for Global AI Weather Forecasting

Abstract

AI weather forecasting models are commonly trained on the ERA5 reanalysis, which is unavailable in real time. Operational deployment therefore relies on initial conditions produced by numerical or AI analysis systems that differ from those encountered during training. This mismatch can degrade forecast skill, while retraining for every analysis system is costly. Here, we present Domain-Adaptive Data Assimilation (DADA), an observation-guided framework that adapts external analyses to pretrained AI weather models. Starting from a background state, DADA optimizes only an initial-state perturbation while keeping the forecast model frozen. The perturbed state is propagated through the model, and its short-range trajectory is constrained by real-world observations through a learned observation operator. The resulting initial condition is shaped jointly by observational constraints and the dynamics learned by the target model. We evaluate DADA across five global AI weather models using backgrounds from the Global Forecast System and the AI-based HealDA. Across deterministic and probabilistic forecasts, DADA substantially reduces short-range skill loss caused by changes in the initial-condition source. More broadly, DADA turns observations into a common interface between independently developed analysis and forecasting systems, enabling pretrained AI weather models to accommodate evolving operational initial conditions without reconstructing ERA5 or retraining the forecast model.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Minseok Seo, Noah Brenowitz, Doyi Kim, Hyesook Lee, Changick Kim. 2026-10-02. Domain-Adaptive Data Assimilation for Global AI Weather Forecasting. https://arxiv.org/abs/2610.03172

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An Integrated Scientific AI Framework for Inorganic Materials Design and Industrial Process Optimization

Artificial Intelligence (AI) is redefining the frontiers of scientific domains, ranging from drug discovery to meteorological modeling, yet its integration within industrial manufacturing remains nascent and fraught with operational challenges. To bridge this gap, we introduce Aethorix v1.0, an AI agent framework designed to overcome key industrial bottlenecks, demonstrating state-of-the-art performance in materials design innovation and process parameter optimization. Our tool is built upon three pillars: a scientific corpus reasoning engine that streamlines knowledge retrieval and validation, a diffusion-based generative model for zero-shot inverse design, and specialized interatomic potentials that enable faster screening with ab initio fidelity. We demonstrate Aethorix's utility through a real-world cement production case study, confirming its capacity for integration into industrial workflows and its role in revolutionizing the design-make-test-analyze loop while ensuring rigorous manufacturing standards are met.

cs.CE↗

RipplePLM: Structural and Property Decoupling for Protein Mutation Effect Generation

Protein mutation effect generation asks a model to describe the functional consequence of a point mutation in natural language. Existing protein-to-text systems typically encode mutation information into undifferentiated representations, overlooking the organization of mutation-induced evidence across structural and biochemical factors. We propose RipplePLM, a mutation-aware generation framework centered on Direct-Distal Cross-Attention (DDCA). By constructing a residue-level Mutation Perturbation Field from pre-trained protein language models, DDCA leverages predicted contact maps to organize mutation representations into two pathways: the mutation site's immediate contact neighborhood and its multi-hop distal context. To complement this structural decomposition, we further introduce the Property Latent Chain (PLChain), which injects expert-guided supervision of biochemical property changes (e.g., thermostability and optimal pH) into the LLM hidden-state pathway through latent property tokens. On MutaDescribe, RipplePLM improves over mutation-specific baselines on temporal and structural splits; under a matched-backbone comparison, average structural-split ROUGE-L increases from 22.23 to 35.65. Expert evaluation further shows a higher proportion of biologically accurate or relevant descriptions than the mutation-specific baseline. Additional ablations, representation diagnostics, and low-$N$ fitness regression experiments further support the effectiveness of the learned mutation-aware representations. Code: https://github.com/Lyu6PosHao/RipplePLM.

cs.CE↗

Data-Free Weak-Form Staggered Neural Operators for Magneto-Mechanical Coupling in Finite-Strain Elastomers

Magneto-active elastomers, as a class of smart materials, exhibit strongly coupled magnetic and mechanical behavior at finite strains. Considering variations in microstructure, material properties, and geometry can lead to computationally expensive analyses. Building on the finite operator learning (FOL) framework, this study develops a data-free, physics-informed operator-learning framework for families of coupled finite-strain magneto-mechanical boundary-value problems. The governing magnetostatic and mechanical equations are enforced through finite-element weak-form residuals, allowing the neural operators to be trained without labeled finite-element solution data. The main contribution of this study is the development of a weak-form staggered neural operator (WSNO) framework for strongly coupled magneto-mechanical saddle-point problems. Separate magnetic and mechanical neural operators are trained alternately using a staggered optimization strategy, while their physical coupling is retained through the constitutive relations and residual evaluations. The resulting framework learns mappings from parameterized material and geometric descriptions to the corresponding coupled magnetic and mechanical solution fields. The proposed framework is investigated across several settings, including heterogeneous random-inclusion microstructures, varying magnetic phase contrast, area-fraction-dependent geometries, strongly out-of-distribution material morphologies, and three-dimensional geometry-parametric problems. In addition, the learned operator is combined with a neural-initialized Newton strategy, in which the nonlinear finite-element solver is initialized using the neural prediction. The results demonstrate that the proposed operator-learning framework can accurately capture coupled magneto-mechanical responses across a broad range of parametric problem settings.

cs.CE↗