Search arXiv⌕ Search

arXiv · 2610.05570

Poisson-GENERIC Neural Operators: Exact Metriplectic Structure in Function Space via Casimir Entropies

Abstract

Existing thermodynamically consistent neural operators impose the GENERIC degeneracy conditions by projecting the reversible operator onto the complement of the entropy gradient. This makes the operator state-dependent and forfeits the Jacobi identity, so the result is metriplectic-degenerate rather than metriplectic. We instead obtain degeneracy the way GENERIC does. For nonlinear transport, the reversible operator $L$ is the compatible Lie-Poisson pencil $αD+λ(uD+Du)$; otherwise it is a constant, trivially Poisson Fourier multiplier. On an augmented state $(u,s)$ with a latent entropy density, $S=\int s$ is a Casimir of $L$, so $L\,δS/δz=0$ holds identically without projection. The energy combines a fixed mechanical quadratic, a learned gauge-free potential, and a convex internal energy. The friction operator $M=AA^\top$ satisfies $M\,δE/δz=0$ pointwise, and its Onsager parity structure permits diffusion and damping while provably excluding transport. For any parameters, skewness, positivity, both degeneracies, and the Jacobi identity (on the resolved band for the Lie-Poisson term) hold to machine precision. Heat conduction and damped waves admit exact closed-form friction operators, the second law bounds physical energy under a checkable curvature condition, and a discrete-gradient integrator yields exact discrete first and second laws. On four PDEs in 1D and 2D with three backbones (FNO, Transolver, CNO), the model wins 61 of 72 seed-level comparisons against same-backbone unconstrained baselines, learns the exact transport and wave symbols, matches the true dissipation rate within 13% on heat and Burgers, and dissipates nothing on advection. A constant-$L$ ablation isolates the cost of exact Jacobi as the loss of Burgers, while a learned-entropy ablation injects energy on every reversible-irreversible problem.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jason Sulskis, Sathya Ravi. 2026-10-04. Poisson-GENERIC Neural Operators: Exact Metriplectic Structure in Function Space via Casimir Entropies. https://arxiv.org/abs/2610.05570

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Modified Loss of Momentum Gradient Descent: Fine-Grained Analysis

We analyze gradient descent with Polyak (1964) heavy-ball momentum (HB) whose fixed momentum hyperparameter $β\in (0, 1)$ provides exponential decay of memory. Building on Kovachki and Stuart (2021), we prove that on an exponentially attractive invariant manifold the algorithm is exactly plain gradient descent with a modified loss, provided that the step size $h$ is small enough. Although the modified loss does not admit a closed-form expression, we describe it up to $O(h^{\mathcal{R}})$-errors for arbitrary finite order $\mathcal{R}$, and prove global (finite "time" horizon) trajectory approximation bounds $O(h^{\mathcal{R}})$. We then conduct a fine-grained analysis of the combinatorics underlying the memoryless approximations of HB, in particular, finding a rich family of polynomials in $β$ hidden inside which include and lie coefficient-wise in between Eulerian and Narayana polynomials. We prove that these polynomials are $h$-polynomials of certain graph-associahedra. As corollaries of the main results, we derive continuous modified equations of arbitrary finite approximation order (with rigorous bounds) and the principal flow that approximates the HB dynamics, generalizing Rosca et al. (2023). Approximation theorems cover both full-batch and mini-batch HB. The results shed new light on the main features of HB and outline a roadmap for similar analysis of other optimization algorithms.

cs.LG↗

HomID : Benchmarking Intrinsic Dimension Estimators on Homogenous Manifolds with Anisotropic Embeddings

The manifold hypothesis suggests that data lies on manifolds with smaller intrinsic dimension (ID) than their ambient dimension. However there is no empirical agreement on the estimates for ID from different estimators for realistic datasets. Thus it is important to test ID estimators (IDEs) with targeted stressors. In this work, we consider the role of anisotropy. To this end, we propose HomID, a collection of homogeneous spaces with anisotropic embedding, for benchmarking ID estimators. We observe that methods that perform well on standard benchmarks systematically degrade on HomID under identical resource allocation. We further observe that anisotropic distortion of such benchmarks also results in performance degradation. Finally, we demonstrate that controlled anisotropic distortions induce systematic shifts in the distributions on which these methods rely, providing a concrete mechanism for the resulting estimation errors in two particular IDEs.

cs.LG↗

Routing-Aware Safety Alignment for Mixture-of-Experts Models

Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.

cs.LG↗